amXor

Blogging about our lives online.

Showing posts with label Assembly. Show all posts
Showing posts with label Assembly. Show all posts

3.03.2010

Assembly Language For Mac

0 comments
I'm away from my Linux box and want to do some assembly programming. Mac installs GCC with the developer tools, but there are enough differences that I haven't bothered to work through them until now. Here's a decent tutorial, although it focuses on PPC assembly and I'm using an Intel Mac. The thing that frightened me about the Mac assembler was the default output of gcc -S. There is some strange optimizations and flags in the resulting assembly code. The key, as the tutorial points out, is in the compiler options. Here's what I used on the ubiquitous "Hello World" program:
gcc -S -fno-PIC -O2 -Wall -o hello.s hello.c
And here's the assembly code it spit out:
    .cstring
LC0:
   .ascii "Hello World!%d\12\0"
   .text
   .align 4,0x90

.globl _main
_main:
   pushl   %ebp
   movl    %esp, %ebp
   subl    $24, %esp
   movl    $12, 4(%esp)
   movl    $LC0, (%esp)
   call    _printf
   xorl    %eax, %eax
   leave
   ret
   .subsections_via_symbols
This is more familiar territory, the only differences being the .cstring directive instead of .section .text, the leading underscore on printf, and the .subsections_via_symbols directive. The general naming of sections is outlined on the Mac Assembler Reference, and the .subsections_via_symbols explanation is interesting. I'm already used to using many labels in my code; does this mean that the named sections would be ripped out because they are not "called" by any other code? I tested this out in the previous example, just adding a second call to _printf in a labelled section and the code worked just fine. It seems that labels don't count, they have to be declared sections like .globl, .section or whatever. That seems fair, I haven't yet made a habit of calling sections that are supposed to flow naturally into other sections. Maybe there is some instance where this might be a useful optimization? I will be looking into Position Independent Code(PIC) a bit more, it seems that it's similar in theory to how the latest Linux kernel runs code at randomized memory locations to prevent hardcoded attacks, but I don't know if that's the extent of it.

3.02.2010

RPN Calculator - v0.02

0 comments
My calculator code was quite easily polished up. Here's the revised code which stacks operands properly and supports the main arithmetic operators +,-,*,/. If you flush the stack completely, you get a "nan" warning, which seems reasonable. Here's the code: (gas, x86)
.section .data
expr_length:    .int 128
ADD:            .ascii "+"
SUB:            .ascii "-"
MUL:           .ascii "*"
DIV:            .ascii "/"
null:           .ascii "\0"
disp_float:     .ascii "%f\n\n\0"
.section .bss
    .lcomm expr, 128
.section .text
.globl main

main:
    finit
    1:
    leal    null, %esi          #Clear the expr buffer
    leal    expr, %edi
    movl    expr_length, %ecx
    cld
    lodsb
    rep     stosb
    addl    $4, %esp
    pushl   stdin               # Read an expression
    pushl   $64
    pushl   $expr
    call    fgets
    addl    $12, %esp

    movb    ADD, %ah            # Test For Operators
    movb    expr, %bh
    cmp     %ah, %bh
    je      addFloat
    movb    SUB, %ah
    cmp     %ah, %bh
    je      subFloat
    movb    MUL, %ah
    cmp     %ah, %bh
    je      mulFloat
    movb    DIV, %ah
    cmp     %ah, %bh
    je      divFloat

    pushl   $expr               # Must be a number
    call    atof
    addl    $4, %esp
    jmp 1b

    addFloat:
        faddp
        fstl   (%esp)
        jmp     disp_answer

    subFloat:
        fsubrp
        fstl   (%esp)
        jmp     disp_answer

    mulFloat:
        fmulp
        fstl    (%esp)
        jmp     disp_answer

    divFloat:
        fdivrp
        fstl   (%esp)

    disp_answer:
        pushl   $disp_float
        call    printf
        addl    $8, %esp
        jmp     1b


    notfound:
    movl $1, %eax
    movl $0, %ebx
    int $0x80

Simple RPN Calculator

0 comments
Reverse Polish Notation is a simple method of calculation that was used extensively in scientific calculators such as the HP 32S but has fallen out of use somewhat these days. I have decided to implement an RPN calculator in Assembly Language to test what I have learned so far. Here is version 0.01 that only adds and subtracts, just to give a rough layout of how it will work.
.section .data
  expr_length:  .int 128
  ADD:          .ascii "+"
  SUB:          .ascii "-"
  null:         .ascii "\0"
  disp_float:   .ascii "%f\n\n\0"
.section .bss
  .lcomm expr, 128

.section .text
.globl main

main:
    finit
1:
    leal    null, %esi          # Clear the expr buffer
    leal    expr, %edi
    movl    expr_length, %ecx
    cld
    lodsb
    rep     stosb               #__
    addl    $4, %esp
    pushl   stdin               # char * fgets (char * str, int len, * stream)
    pushl   $64
    pushl   $expr
    call    fgets
    addl    $12, %esp

    movb    ADD, %ah            # Test For Operators in expr
    movb    expr, %bh           # and jump to code if operator found
    cmp     %ah, %bh
    je      addFloat
    movb    SUB, %ah
    cmp     %ah, %bh
    je      subFloat

    pushl   $expr               # Defaults to a number
    call    atof                # double atof (const char * str)
    addl    $4, %esp            # pushes a float into st(0) from string
    jmp 1b


addFloat:
    faddp
    fstpl   (%esp)
    pushl   $disp_float
    call    printf
    addl    $8, %esp
    jmp     1b

subFloat:
    fsubrp
    fstpl   (%esp)
    pushl   $disp_float
    call    printf
    addl    $8, %esp
    jmp     1b

    movl $1, %eax           # All roads jump back to 1 so we never get here
    movl $0, %ebx
    int $0x80
And here's how it looks in action:
$./rpn
14.5
88
+
102.500000

12.3
6.6
-
5.700000
It took some digging to figure out how to load a value into the FPU from a string and the "atof" C library function seemed to be the easiest. All it needs is a string pointer on the stack ("$expr" in the code) and it does it's best to convert it to a float and push it onto the FPU stack - fairly painless. The operator test needs some work since it will always give a false positive if you punch in a signed value. Punching in -12 causes the calculator to run the subtraction code and disregards the number you typed in, not ideal. Other todos: 1. Pull the display code out of the calculations, this should be generic. 2. Push the result back onto the stack so you can use it in following calculations. 3. Filter the input somewhat.

2.18.2010

Assembly Language - Command line parameters

0 comments

I ran into a problem with the way that command line parameters are passed in Assembly Language. The book I'm working through has a couple example programs illustrating how to access the command line arguments and none of these examples worked on my system. The simple solution was that, at some point between this book being written and me reading it, this convention had changed. Instead of the stack looking like it does in the book, with all the command line arguments on the stack, we have the number of parameters first and then a pointer to a pointer to all the parameters in null terminated strings.

This is how it looks in the old style:
number: 3
*p1: program name
*p2: param1
*p2: param2

And this is how it looks in the new style:
number: 3
*p: *params
*params: param1, param2, ...

I should have guessed this from the C calling convention **args, but it took a fair bit of digging with GDB to analyze what exactly was going on with the stack. So, the end result: A program to print all the command line arguments:

.section .data
output:
    .asciz "parameter: %s\n"
.section .bss
.section .text
.globl main

main:
    nop
    movl    (%esp), %ebx
    movl    4(%esp), %ecx      # Num parameters
    movl    8(%esp), %ebx      # This is a pointer to the string pointer
    movl    (%ebx), %edi       # This should be the string pointer
    
    loopargs:
        pushl   %ecx
        pushl   %edi
        pushl   $output
        call    printf
        addl    $8, %esp
        movl    $0x255, %ecx
        movb    $0, %al 
        cld
        repne   scasb
        popl    %ecx
        loop    loopargs
        
    nop 
    movl $1, %eax
    movl $0, %ebx
    int $0x80

2.10.2010

px Language

0 comments
In my studies with assembly language, a couple of things have occurred to me: 1. Assembly syntax is SIMPLE. 2. Instructions, registers, stack pointers, etc. are incredibly confusing. What i mean by the syntax being simple, is that it is completely linear. If the code jumps around it explicitly tells you where it's jumping and why. All lines can be understood as self-contained statements and the only structure to the program is the structure you give it. With this in mind, I have set out to design a new language based on this syntax. The language will be functional and type-safe like Haskell, but it won't have ANY built in syntax. I have begun to build a parser in Python which does the px to asm conversion. It's very simple at this stage, but should work. The language is read, for the most part, right to left with the leftmost term(s) being the terminal element. The only exception to the rule might be infix operators, because I'm not sure if many would take up a language if they had to learn reverse-polish notation. The Basics Here's a simple example of naming symbols and comparing them.
x, y  20, 15
z 42
z x if x > y
z y if y > z
print z
x, y and z are assigned the values 20, 15 and 42. y and x are pushed onto the stack and '>' pops them and compares them, putting a 0 or 1 on the stack. The if statement takes two operands, the x and the boolean result of the '>' operation. It then pushes onto the stack x and [0/1] (secret rule #1). Secret rule#1: Values are always stored in two parts, a value and a flag. If the flag is false the assignment doesn't happen. I have chose to omit the assignment operator, because all evaluations logically flow right to left. You will see the benefit of using infix operators if you reduce this operation to it's RPN form, with parens added for clarity: z (if ( x (> y x))) The > and other mathematical terms can be understood with some practice in RPN, but the if statement is not very intuitive at all without infix notation. Another curiosity that needs addressing is that since this is stack-based assignment, the first assignment "x,y 20,15" just adds the numbers 15 and 20 onto the stack and then pops them off making y=20 and x = 15. Not the preferred functionality! Functions Simple enough for basic expressions, how do we implement more complex functions? I've carried the same logic into function expressions. Functions always produce the left-hand value based on the right-hand input, but you add a code block below the statement line. This is how it looks in practice:
x 15
y 12
.x some_function .x, .y:
   .x sum(.x, .y)
   :
print x
I almost went with the Python-esque whitespace-significance, but have decided for now to close blocks with a final colon. The function's header defines it's inputs and outputs. The function takes two values, x and y, passes them through the function block and produces x. There are no return statements, because the output symbol is explicitly given in the header line. Okay, the syntax is simple enough, but what can you do with it, and what are the leading periods all about? That's where namespaces come into play... Namespaces Each block defines it's own namespace. The previous function declared inputs of .x and .y. These are concrete references to the parent namespace. The output was also a concrete reference, so the function was fully defined and executed to produce the new value of x. Here's an example that, although it looks very similar, is actually quite different:
x 15
y 12
m generic_function m, n:
  m sum(m, n)
  :
x generic_function x, y
print x
This is more like normal function you would use, because it defines the function with generic inputs and outputs and then you use them on specific values. The function is only executed when called with some specific values. This opens the door for partially applied functions and the like. eg:
m part_func .x, m:
   m sum(.x, m)
   :
y part_func y

m const_function .x, .y:
  print x, y
  m True
  :
res const_function

.x defined_out m, n:
   .x sum(m, n)
   :
defined_out 8 1024
print x
This is a simple but powerful way to implement a lot of the functionality of high level language with only one syntactic convention. One final note about namespaces is that the .* symbol references the entire parent namespace. This allows one to implement transparent code blocks like loops. Eg:
x 0
x_max 11
.* loop_block .*:
   continue .x < .x_max     print .x     .x .x + 1     loop True     : 
Because the input and output are fully defined, the code executes and can access any of the variables, functions etc. of the parent block via the dot syntax. The loop and continue functions control the flow of the block and can be used in any block. 'loop' is just a "goto $start if expression is true" statement and 'continue' is a "goto $end if expression is false" statement. This block first checks it's condition, breaking if it is false, does some more stuff and then loops back to the start unconditionally. Alternately, one might want to omit the 'continue' statement and do the condition check with the 'loop' statement. Every block will have a start and end tag that can be used in this manner. Objects Because blocks are self-contained namespaces, there is another possible use for them, classes and objects.
. some_object .:
   name "some_object"
   value 42
   :
This block didn't specify any inputs or outputs. It is thus a concrete object which can be used as follows:
print some_object.name
But maybe I'm getting ahead of myself! Stay tuned for more developments and please give suggestions on the name, I'm not sure how or why I came up with "px".

2.09.2010

Assembly Language - Development Environment

0 comments
The development environment for assembly language outlined in this book is a bare bones, linux-based toolchain. There is no one-click "Build and Go" or "Build and Debug" options here. I have installed the tools on my Arch Linux backup system which I am using through ssh from my laptop.
Toolchain
Assembler: GAS (GNU Assembler, part of the GCC Package).
Linker: ld, included with gcc.
Debugger: gdb (wasn't installed, `pacman -S gdb` solved that).
Profiling Tools: objdump, gprof are included in the binutils package.
I also looked into setting up a development environment using Xcode on my mac. It is fairly trivial to set up a "Standard Tool" project and call assembly code, but the assembly code uses some different conventions on mac vs. Linux. I've decided to stick with the linux style used in the book, but there is a good overview of the basics of Mac assembler in the Mac developer reference.
Using The Assembler
So, a couple notes about using the tools. My first observation was that to know what a program is doing in assembly, you must learn to use the debugger, always compile with -gstabs. The book recommends two stage compilation, assemble and link, but I'm not sure why. A one stage gcc invocation guarantees(?) proper linking without having to explicitly name the library and the dynamic linker.
Example: Using as & ld when linking to a C library
as -o somefile.o somefile.s
ld -dynamic-linker /lib/ld-linux.so.2 -o somefile -lc somefile.o
Compared to using gcc:
gcc -o somefile somefile.s
I think I'll stick with gcc for now. And including debugging symbols is straightforward as well:
gcc -gstabs -o somefile somefile.s
Using the Debugger
The debugger is not quite as mystical when used in assembly language. For the most part each line corresponds to a machine instruction, so you can follow along in the source code and view the results of each operation. When the debugger starts, it prepares the code for execution and then awaits your instructions. If you simply type 'run' it will execute the code to the end. To see what the program is doing, you must specify a breakpoint: 'break *main' stops execution at the start of the main block. From here we can step through the program with the 'step' instruction. Some interesting things to be looking at during execution are the registers 'info registers', data values 'x/d &value' and the stack 'info stack'.
Compiling C source to Assembly
Another good resource for studying assembly language is to write the program in C and compile it with the -S option which outputs the assembly equivalent of the C code. This may also be a good way to fine-tune your C code, but for the purposes at hand it's simply a learning exercise, and I don't plan on doing better than GCC any time soon!
Source Text: Professional Assembly Language (Blum, 2005)

2.05.2010

Assembly Language - Instruction Codes

0 comments
Source Text: Professional Assembly Language (Blum, 2005) Chapter 1:
This book begins with a brief overview of CPU instruction code handling, which is still a bit of a mystery after reading it but I will try to describe it to the best of my abilities. Here's a quick sketch of how the CPU processes instructions:
Instruction code <- Instruction Pointer <- Memory 
   Data <- Instruction code (Data Element, opt.)
   Data <- Registers
   Data <- Data pointer <- Memory 
Each instruction code can perform operations on data from the registers, main memory, or it's own 1-4 byte data elements. The actual instruction code can be from 1 - 17 bytes long, the only required byte(s) being the opcode which tells the processor what operation to perform. The other bytes of the instruction code specify modifiers for the operation and data elements (0-4 bytes) for the operation. Data registers are essentially memory locations on the processor itself. This makes them the fastest way to store bits of information, but the space is very limited (x86 defines eight 32bit registers). Assembly language adds a layer of abstraction to this by adding mnemonics for the opcodes.
55
89 E5
83 EC 08
Translates to:
 push %ebp
mov %esp, %ebp
sub $0x8, %esp
It also lets you declare and name your data.
testvalue:
 .long 150
message:
 .ascii "This is a test message."
pi:
 .float 3.14159
In summary, assembly language may look extremely cryptic, but it is at least human-readable and it gives the programmer access to the core of a CPU's processing. So here's my first assembly language program:
# cpuid.s Sample program to extract the processor Vendor ID
.section .data
output:
.ascii "The processor Vendor ID is 'xxxxxxxxxxxx'\n"

.section .bss

.section .text
.globl _start

_start:
movl $0, %eax
cpuid
movl $output, %edi
movl %ebx, 28(%edi)
movl %edx, 32(%edi)
movl %ecx, 36(%edi)
movl $4, %eax
movl $1, %ebx
movl $output, %ecx
movl $42, %edx
int $0x80
movl $1, %eax
movl $0, %ebx
int $0x80
Which I assembled and linked with the following commands:
$ as -o cpuid.o cpuid.s
$ ld -o cpuid cpuid.o
And the output:
$ ./cpuid
The processor Vendor ID is 'AuthenticAMD'
Hooray!

Twitter

Labels

Programming (21) Assembly (7) Productivity (5) Goals (3) Planning (3) Writing (3) cognition (1)

Followers

andyvanee.com

Files