
Okay, this one isn't included with your Mac, but it really should be. You have to download Megazoomer as well as SIMBL.
Okay, this one isn't included with your Mac, but it really should be. You have to download Megazoomer as well as SIMBL..section .data
expr_length: .int 128
ADD: .ascii "+"
SUB: .ascii "-"
null: .ascii "\0"
disp_float: .ascii "%f\n\n\0"
.section .bss
.lcomm expr, 128
.section .text
.globl main
main:
finit
1:
leal null, %esi # Clear the expr buffer
leal expr, %edi
movl expr_length, %ecx
cld
lodsb
rep stosb #__
addl $4, %esp
pushl stdin # char * fgets (char * str, int len, * stream)
pushl $64
pushl $expr
call fgets
addl $12, %esp
movb ADD, %ah # Test For Operators in expr
movb expr, %bh # and jump to code if operator found
cmp %ah, %bh
je addFloat
movb SUB, %ah
cmp %ah, %bh
je subFloat
pushl $expr # Defaults to a number
call atof # double atof (const char * str)
addl $4, %esp # pushes a float into st(0) from string
jmp 1b
addFloat:
faddp
fstpl (%esp)
pushl $disp_float
call printf
addl $8, %esp
jmp 1b
subFloat:
fsubrp
fstpl (%esp)
pushl $disp_float
call printf
addl $8, %esp
jmp 1b
movl $1, %eax # All roads jump back to 1 so we never get here
movl $0, %ebx
int $0x80And here's how it looks in action:$./rpn 14.5 88 + 102.500000 12.3 6.6 - 5.700000It took some digging to figure out how to load a value into the FPU from a string and the "atof" C library function seemed to be the easiest. All it needs is a string pointer on the stack ("$expr" in the code) and it does it's best to convert it to a float and push it onto the FPU stack - fairly painless. The operator test needs some work since it will always give a false positive if you punch in a signed value. Punching in -12 causes the calculator to run the subtraction code and disregards the number you typed in, not ideal. Other todos: 1. Pull the display code out of the calculations, this should be generic. 2. Push the result back onto the stack so you can use it in following calculations. 3. Filter the input somewhat.
I ran into a problem with the way that command line parameters are passed in Assembly Language. The book I'm working through has a couple example programs illustrating how to access the command line arguments and none of these examples worked on my system. The simple solution was that, at some point between this book being written and me reading it, this convention had changed. Instead of the stack looking like it does in the book, with all the command line arguments on the stack, we have the number of parameters first and then a pointer to a pointer to all the parameters in null terminated strings.
This is how it looks in the old style: number: 3 *p1: program name *p2: param1 *p2: param2 And this is how it looks in the new style: number: 3 *p: *params *params: param1, param2, ...
I should have guessed this from the C calling convention **args, but it took a fair bit of digging with GDB to analyze what exactly was going on with the stack. So, the end result: A program to print all the command line arguments:
.section .data
output:
.asciz "parameter: %s\n"
.section .bss
.section .text
.globl main
main:
nop
movl (%esp), %ebx
movl 4(%esp), %ecx # Num parameters
movl 8(%esp), %ebx # This is a pointer to the string pointer
movl (%ebx), %edi # This should be the string pointer
loopargs:
pushl %ecx
pushl %edi
pushl $output
call printf
addl $8, %esp
movl $0x255, %ecx
movb $0, %al
cld
repne scasb
popl %ecx
loop loopargs
nop
movl $1, %eax
movl $0, %ebx
int $0x80
In order to more fully understand how I'm going to implement the translator from my imaginary programming language, I have written a few direct translations from px to assembly. My thinking is that if the language doesn't rely on the abstractions provided by the "compile to C" interpreted languages, I should be able to end up with a systems level language with the clean syntax of python, ruby, et al.
My first example is to calculate the sum of a large list. Here are two versions in python.
UPDATE!!! After trying a few more versions, I realized that my loop counters were off in every example. I had made the silly python mistake of forgetting the range of range. range(0, 123456789) = 0,1,2...123456788. And I had subsequently propagated that mistake into my C and Assembly versions, producing the wrong answer in every case, sometimes very fast! Lesson learned: Check your math!
# listsum1.py - calculate 1 + 2 + ... + 12,345,678 = listsum
listsum = 0
for f in range(0,12345678):
listsum = listsum + f
print "The result is " + str(listsum)
# listsum2.py
listsum = sum(range(0,12345678))
print "The result is " + str(listsum)And here's what it might look like in px.
print listsum buildlist 12345678
But, of course, no library functions exist yet, so I'll spell out how the functions listsum and buildlist would be defined.
n... buildlist size, next 0:
next next + 1
n next if next < size
:
n listsum list...:
break if not list
n n + list
continue True
:buildlist is defined as taking one argument, size and producing a list, one item at a time. Listsum consumes these items and produces a final value when the whole list has been read. All items are passed as two 32-bit values, [type/status] and [value/pointer]. Each operation checks the status flag to check if it is receiving a valid object and what that object is exactly. The buildlist function uses this fact to push a false object onto the stack when it reaches "size". This tells listsum that there are no more values to be read. In assembly this is what the code structure looks like in simplified form:
list_size: .int 12345678
l_sum: .int 0
push $1
push list_size
pop list_size
pop edi #using edi as the status flag
mov 0, ecx #and ecx for the counter
buildlist:
compare list_size, ecx
cmovae 0, edi #if above, set edi to 0 (false)
add 1, ecx
push edi
push ecx
listsum:
pop eax
pop edi
cmp 0, edi
je end_listsum
add eax, l_sum
jmp buildlist
end_listsum
print l_sumIn the actual implementation, I changed the add to support a quadword because the results get large quite fast, but this gives you the idea.
Well, I knew it would be fast, but this is a pretty large list, the code between buildlist and end_listsum has to run over 12 million times to calculate the sum, this is what I got when comparing the two python versions and my assembly version:
$ time python listsum1.py The result is 76207876467003 real 0m15.379s user 0m14.819s sys 0m0.557s $ time python listsum2.py The result is 76207876467003 real 0m8.123s user 0m7.476s sys 0m0.647s $ time ./listsum The result is 76207876467003 real 0m0.275s user 0m0.273s sys 0m0.000s
That's pretty blazing fast if you ask me! But I was put in my place when I ran the C version which produced these results:
time ./f_sum3 The answer is 76207876467003.000000 real 0m0.123s user 0m0.120s sys 0m0.000s
I used a floating point double for the list sum, so maybe that's why C was faster. Still, if we use C as the baseline we have:
C 1 px 2.2 Py2 66.04 Py1 125.03
Not bad for a start. Now I'm off to dig through the compiled C code to see how it should be done...
UPDATE #2 - I've realized that there is a quite simple algorithm to solve this particular problem and that is: (n/2) * (n+1). This simple equation eliminates the loop completely making the calculation trivial. The updated python version runs in 0.038s and the C version runs in 0.002s producing the, now correct sum, 76207888812681.
After spending most of the day investigating parsers, lexers and compiler compilers, I've come to the conclusion that my language doesn't really need these just yet. I've really only outlined three syntactic rules and two structural ones, and these cover all the ground I can see in front of me.
z y if (y > x) z x if (y2. Right to left definition of block header line:
z function z, y: z do_something y :3. Comma seperated values are globbed together into an array.
process_list x, y, 12, 72, "Hello World!"
That's basically it for the syntax, where it gets interesting is namespaces and typing.
Every code block defines it's own namespace by default. This means that every block of code is explicit in defining what values it is using. Here's a trivial block which takes two operands and produces another:
.z a_function .z, y: .z y + .z
The ".", and the lack thereof, are the essential parts when it comes to namespaces. You might have seen the dot syntax in other languages (my_array.append()) and it's not really that different in px, except that referencing variables is a two way street. a_function is declaring it's own namespace under the main program and thus becomes main.a_function. That's fine, everything in the main namespace can access it directly by it's bare name and we can keep stacking them just like in other languages.
But the trick here is the preceding '.' on the variable name z and the lack of it on y. What's actually going on here? What happens is that the function is defined with concrete references to the z in the parent namespace and one reference to a y which does not exist. This becomes a unary function, a partially applied function and is never called until it is supplied with a y. When it does get called with a value, the z in the parent namespace gets updated accordingly. This function might be more appropriately called "increment_z", let's see it in action:
.z 10 increment_z 1 print z >>11 increment_z 8 print z >>19
A standard for loop imports all names, but it does so explicitly in it's function definition. "Hold on," you might ask, "it's function definition???" How does a for loop have a function definition? Don't other languages define this as built in to the language? The answer is yes, but in px I have decided to expose all the gory details of language design to the programmer, and only hardcode the bare essentials. This is roughly how the standard loops are implemented as blocks:
.* loop .*, _start, _end: break _end condition continue _start condition :
The _start and _end tags are handles for the final assembly language implementation. The * wildcard states that any and all of the parent namespace will be accessible. This block definition has no pre-determined outputs or effects and so it's considered a function template. If called, the break and continue expressions will only be included if their conditions are met. Here's how that might look:
loop: break if .x
Loop doesn't actually control the looping, the statements within it do. We could have omitted the break statement and used the condition of the continue to control the loop. Notice again that this block is a fully defined namespace and all external references must be used accordingly.
To summarize namespaces, block = namespace = object = method = class = control mechanism = symbol. If a block has fully defined inputs and outputs, it is treated just like any other literal value in the flow.
I've stated that there are two structural elements to my language: Namespaces and Types. Well, I kind of lied: there's only one types are also implemented functionally, but I haven't worked out how and how much of the checking can be done in the final assembly code. The general idea is this: values that aren't known at compile time are put on the heap and indexed with a hash table. Values that are know (by single literal assignment, or explicit typecasting) are put on the stack / data segment. Type statements are just like any unary operator:
x number 225.993 y int 235 z string "Please assign me to z!"
They go right next to the value being assigned as a check for the correctness of the data being passed to it.