. ~ . ~ ~~ . ~~ ~~~ ~ ~~~ ~~~~ ~ ~~~~~~~ ~~~~~ ~~~~~~~~~~~~~~ . ~~ Evocation! ~~ ~~~~~~~~~~~~~~~ or, how to call the blue-green flame ~~~~~~~~~~~~~~ ~~~~~~~~~~~~~~~ The documentation is a work in progress. It doesn't say most of the things it needs to, yet. Evocation is a dialect of Forth, grown to Irenes' tastes. It is meant to someday be a platform for experimenting with parse theory, type theory, databases, and other things Forth is not traditionally known for, as well as with language design, which it is. It is a self-hosting compiler, meaning the only thing you need to build it is a copy of itself. At present, Evocation targets only one architecture, amd64. It is rare among compiled Forths in that it targets a 64-bit architecture. Someday, Evocation will also be self-bootstrapping, meaning that it will be able to "compile" itself into a commented hex dump of itself for ease of auditing. This rests on the insight, from the mescc and guix developers, that the difference between source code and binary is comments. The efforts in this direction are described below under "Hexing Evocation for Distribution". ~~~~~~~~~~~~ ~~ Building ~~ ~~~~~~~~~~~~ Since we have chosen not to distribute Evocation in binary form, you don't have a copy of it yet and cannot take advantage of its self-hosting properties for your first-ever version. Happily, until the self-bootstrapping properties are ready, we have maintained compatibility with the original version of Evocation which was written in a program called flatassembler, which you will have to acquire. TODO mention the terminal To get started, first build the flatassembler version: $ fasmg quine.asm quine $ chmod 755 quine It's called "quine" for historical, sentimental reasons having to do with the original architecture. It is not a Quine, in the sense that it is not a program that outputs its own source. When the self-bootstrapping is done, it will be moved to a historical subdirectory and this misnomer will no longer matter. This "quine" binary is a working Evocation interpreter, but it's incomplete and will become more so with time. So, next, build Evocation-in-Evocation: $ (cat labels.e elf.e transform.e execution.e; echo 's" pyrzqxgl" allocate-string dup 262144 read-to-buffer'; cat core.e linux.e output.e amd64.e execution-support.e log-load.e; echo pyrzqxgl swap 262144 read-to-buffer; cat core.e linux.e output.e amd64.e execution-support.e log-load.e dynamic.e input.e interpret.e flow-control.e linux-dynamic.e ; echo pyrzqxgl; cat evoke.e) | ./quine > evoke $ chmod 755 evoke Now keep your "evoke" binary somewhere safe, and use it to build new versions as you modify Evocation. ~~~~~~~~~~~~~ ~~ Exploring ~~ ~~~~~~~~~~~~~ You can now try out Evocation. Type to it interactively: $ ./evoke ." Hi, Irenes!" 6 7 * . newline bye See what it prints! TODO give examples of RPN for arithmetic Some helpful words to try to get started are list-dictionary and describe. TODO show how to use them The syntax for a string literal is s" ...". There's something very subtle happening: it's the lowercase letter "s", a double quote, and a space. Then you type the actual contents of the string, then at the end you type another double quote. The words "Hi, Irenes!", in the example above, are part of a closely related syntax that uses a period instead of a letter s, and prints the string out to your terminal instead of returning it. Whether you're using s" or .", that space after the quote is mandatory, which may seem very strange if you're more familiar with pretty much any language that isn't a Forth, but it's a common Forth idiom. Requiring the space allows the string literal syntax to be tokenized just like any other space-delimited word. Unlike most languages, there's no special concept of an operator or punctuation character that can "interrupt" another word or run up against the start of one. Everything is separated by spaces. Of course, in modern Forth dialects it's also very common to add a special lexer feature for that sort of thing. Evocation doesn't do that, because eventually fancy syntax and grammar will be implemented at a higher layer, using the experimental parsing formalism that doesn't exist yet. If you want to see where this feature would be if it were implemented in the simpler way, you can read the definition of the word named "word", in interpret.e. If you went to look at that and you're wondering: Yeah, Evocation's lexer really is that short and simple. Part of why it's able to be that easy is that words such as s" that introduce special syntaxes do their own lexing for whatever comes after them. TODO show how to define words Evocation has high-level flow-control words: if, unless, if-else, forever, and while. High-level flow control is a common thing for modern Forth dialects to add, but every dialect does it a bit differently. Evocation's flow-control words are postfix operations and work with curly braces, like this: $ ./evoke : count 10 0 { 2dup < } { space dup . 1+ } while 2drop newline ; count What will it print? :) Evocation's high-level flow control works only in compiled code; this example defines and compiles a new word called "count", in order to show it off. If you try to use the "{ ... } { ... } while" syntax outside of a word definition, it won't do what you expect. This is because, unlike some modern Forths, Evocation doesn't have a general-purpose memory management facility; it uses something called the log, which makes it easy to allocate things but hard to deallocate them. In order to loop through a code block, it has to be allocated somewhere. So, the design takes care not to encourage programming habits that would burn through memory space. ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ~~ Reading Evocation's source code ~~ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Even if you're only interested in using Evocation, not in modifying it, we encourage you to at least skim through the source. If you've looked at it, even a little, it won't be so scary next time. It's heavily commented and meant for anyone with a little programming knowledge to be able to read, even if you've never done systems programming before. If you find something in it confusing, please don't be afraid to ask! It's likely other people are confused too, and sharing your questions helps improve the documentation and lets others learn by watching. The top-level source file whose job is to compile Evocation itself is evoke.e. It's really short, and worth a quick glance right now. It lists all the other source files and the order they get loaded in and how they're processed. The files that do the work to make Evocation run at all are execution.e and core.e. It's worth reading through both of them slowly. After you've read core.e, you'll know a lot of basic words that can be used as commands within Evocation. A lot of Evocation is written in Evocation's version of assembly language. The file amd64.e is the one that implements all the assembly instructions. Writing a real program in assembly also requires resolving labels, which are a special syntax that gives names to addresses. The behavior of labels is all implemented in labels.e. On the assembly language front, there's also linux.e which contains assembly words for doing things specific to the Linux operating system, such as reading input, and there's elf.e which contains words for outputting the special file headers that let the operating system understand that a file is an executable program. In terms of the Forth-y bits, input.e and output.e are concerned with getting text into and out of the language. The infrastructure to define words is in dynamic.e, and the syntax for it is in interpret.e. The high-level flow control words are in flow-control.e. Some of the features of execution.e had to be separated out into their own file, because of details about how the compiler works; that stuff is in execution-suport.e. So, there's all those relatively normal compiler internals in those various files, which are all fairly self-contained... and then there's the transformation facility. This is Evocation's most unique archictural decision, and it's in transform.e. It's well documented, but it's also extremely conceptually dense. Feel free to give it a skim, that's the only way to build familiarity with these things, but you should probably have a solid understanding of the rest of the internals before you place any high expectations on yourself around understanding the transformation facility. It's okay, you can benefit from it before you understand it: Transformation provides the core tricks that make it possible to compile Forth code into standalone executables. The call to label-transform in evoke.e, and the call to log-load-transform in execution.e, are the two spots where compilation is handed off to the transformation facility, and you can pretty much just take it for granted that it works, until you feel ready. If you want examples of programs that are smaller than Evocation itself, quine.e is a tiny program written in proper Evocation that outputs its own source code; hello.e is a hello-world written in Evocation-assembly, and hex.e is another small Evocation-assembly program that might make a good example of how to do slightly more complex things that way. All three of these are self-contained, consisting of just that one file plus calls to Evocation's built-in library. ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ~~ Modifying Evocation's Internals ~~ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ There may come a point in your explorations when you wish to make changes to the compiler. When doing so, please always make sure to build Evocation-in-Evocation both via Evocation-from-flatassembler, as described in "Building" above, and with itself, like this: $ (cat labels.e elf.e transform.e execution.e; echo 's" pyrzqxgl" allocate-string dup 262144 read-to-buffer'; cat core.e linux.e output.e amd64.e execution-support.e log-load.e; echo pyrzqxgl swap 262144 read-to-buffer; cat core.e linux.e output.e amd64.e execution-support.e log-load.e dynamic.e input.e interpret.e flow-control.e linux-dynamic.e ; echo pyrzqxgl; cat evoke.e) | ./evoke > evoke2 $ chmod 755 evoke2 The two versions evoke and evoke2 should be bytewise identical; if they are not, please fix that. This is an important property which would be very difficult to get back if we ever lose it, it's easier to maintain it in-the-moment. ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ~~ Hexing Evocation for Distribution ~~ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The long-term strategy for Evocation's binary bootstrapping is not yet ready, but it's described here anyway because this is where the explanation should eventually go, and it's easier to write about the pieces as they're created. The binary bootstrapping strategy rests on something called the "hex transform", the most complex of the transformations provided as part of Evocation's transformation facility in transform.e. The hex transform has the task of transforming an entire compilation process, which would otherwise produce an executable binary, and instead output a commented hex dump of that binary which describes its internals and their purpose, byte by byte, in sufficient detail to allow a human reader to audit their correctnes. It will do this by passing through comments and call-stack information from the compilation process to the resulting output. In order to turn this commented hex dump into a binary, there is a tiny program called "hex" which handles comments in Evocations ~ syntax, and converts ASCII hexadecimal to raw binary. This program is in hex.e and is written in Evocation-assembly. When compiled it is only 480 bytes, which is small enough to fully audit in its raw, binary form. This is slightly larger than necessary; many of those bytes are used for error message strings, on the principle that it's very important that it be easy to distinguish a successful invocation of "hex" from a failed one. Although the hex transform is not yet fully the compiled "hex" has proven quite stable, and the hex transform does work on it. So, a copy of the compiled "hex" is checked into source control so that it can serve as a root of trust for future Evocation builds. For ease of auditing, a commented hex dump version of this binary, produced via the hex transform, is also checked in, as "hex.hex" (We heard you liked metacircularity, so we put some metacircularity in your metacircularity so you can be metacircular while you're metacircular.) If you need to compile "hex", you can do so as follows: $ cat labels.e elf.e hex.e | ./evoke > hex $ chmod 755 hex To produce the hex-dump version of it, do: $ (cat labels.e elf.e transform.e; echo 's" xyzzy" allocate-string dup 262144 read-to-buffer'; cat core.e linux.e output.e amd64.e execution-support.e log-load.e dynamic.e input.e interpret.e flow-control.e linux-dynamic.e labels.e elf.e hex.e; echo 'xyzzy s" hex-source" variable 1024 1024 * allocate s" hex-binary" variable 1024 1024 * allocate s" hex-metadata" variable hex-metadata hex-binary dup hex-source 5 roll hex-transform bye ' ) | ./evoke > hex.hex Although the hex transform doesn't yet work on Forth programs (only programs written in Evocation-assembly), if you intend to play around with this you may wish to know how to attempt to run it on things. The latest draft way to do that is: $ (cat labels.e elf.e transform.e; echo 's" xyzzy" allocate-string dup 1048576 read-to-buffer'; cat core.e linux.e output.e amd64.e execution-support.e log-load.e dynamic.e input.e interpret.e flow-control.e linux-dynamic.e labels.e elf.e transform.e execution.e; echo 's" pyrzqxgl" allocate-string dup 262144 read-to-buffer '; cat core.e linux.e output.e amd64.e execution-support.e log-load.e; echo pyrzqxgl swap 262144 read-to-buffer; cat core.e linux.e output.e amd64.e execution-support.e log-load.e dynamic.e input.e interpret.e flow-control.e linux-dynamic.e; echo pyrzqxgl; cat evoke.e; echo 'xyzzy s" evoke-source" variable 1024 1024 2 * * allocate s" evoke-binary" variable 1024 1024 * allocate s" evoke-metadata" variable evoke-metadata evoke-binary dup evoke-source 5 roll hex-transform bye ' ) | ./evoke > evoke.hex It should run to completion, producing output. The output is even correct in the sense that passing evoke.hex through ./hex will give a binary that's byte-for-byte identical to evoke, but the output has various problems such as displaying assembly parameters in the wrong order, having insufficient explanation of label references and definitions, not showing dictionary entry headers in any special way, and so on. All these cosmetic issues should be fixable now, and should likely be the focus of any development efforts. It also takes several minutes to run, which Irenes believe is because of the use of linked lists rather than hash tables for the various dictionaries. Adding a hash table is a task to do after bootstrapping is complete. Nearly all of this runtime is attributable to the log-load transform; if you're working on something that doesn't involve the log-load transform, you may find it useful to temporarily comment out the call to log-load-transform in execution.e, replacing it with two invocations of drop. Now get debugging! :)