1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
|
. ~ .
~ ~~ .
~~ ~~~
~ ~~~ ~~~~ ~
~~~~~~~ ~~~~~
~~~~~~~~~~~~~~ .
~~ Evocation! ~~
~~~~~~~~~~~~~~~ or, how to call the blue-green flame
~~~~~~~~~~~~~~
~~~~~~~~~~~~~~~
The documentation is a work in progress. It doesn't say most of the things
it needs to, yet.
Evocation is a dialect of Forth, grown to Irenes' tastes. It is meant to
someday be a platform for experimenting with parse theory, type theory,
databases, and other things Forth is not traditionally known for, as well as
with language design, which it is. It is a self-hosting compiler, meaning the
only thing you need to build it is a copy of itself.
At present, Evocation targets only one architecture, amd64. It is rare among
compiled Forths in that it targets a 64-bit architecture.
Someday, Evocation will also be self-bootstrapping, meaning that it will be
able to "compile" itself into a commented hex dump of itself for ease of
auditing. This rests on the insight, from the mescc and guix developers, that
the difference between source code and binary is comments. The efforts in this
direction are described below under "Hexing Evocation for Distribution".
~~~~~~~~~~~~
~~ Building ~~
~~~~~~~~~~~~
Since we have chosen not to distribute Evocation in binary form, you don't
have a copy of it yet and cannot take advantage of its self-hosting properties
for your first-ever version. Happily, until the self-bootstrapping properties
are ready, we have maintained compatibility with the original version of
Evocation which was written in a program called flatassembler, which you will
have to acquire.
TODO mention the terminal
To get started, first build the flatassembler version:
$ fasmg quine.asm quine
$ chmod 755 quine
It's called "quine" for historical, sentimental reasons having to do with
the original architecture. It is not a Quine, in the sense that it is not a
program that outputs its own source. When the self-bootstrapping is done, it
will be moved to a historical subdirectory and this misnomer will no longer
matter.
This "quine" binary is a working Evocation interpreter, but it's incomplete
and will become more so with time. So, next, build Evocation-in-Evocation:
$ (cat labels.e elf.e transform.e execution.e; echo 's" pyrzqxgl" allocate-string dup 262144 read-to-buffer'; cat core.e linux.e output.e amd64.e execution-support.e log-load.e; echo pyrzqxgl swap 262144 read-to-buffer; cat core.e linux.e output.e amd64.e execution-support.e log-load.e dynamic.e input.e interpret.e flow-control.e linux-dynamic.e ; echo pyrzqxgl; cat evoke.e) | ./quine > evoke
$ chmod 755 evoke
Now keep your "evoke" binary somewhere safe, and use it to build new
versions as you modify Evocation.
~~~~~~~~~~~~~
~~ Exploring ~~
~~~~~~~~~~~~~
You can now try out Evocation. Type to it interactively:
$ ./evoke
." Hi, Irenes!"
6 7 * . newline
bye
See what it prints!
TODO give examples of RPN for arithmetic
Some helpful words to try to get started are list-dictionary and describe.
TODO show how to use them
The syntax for a string literal is s" ...". There's something very subtle
happening: it's the lowercase letter "s", a double quote, and a space. Then
you type the actual contents of the string, then at the end you type another
double quote. The words "Hi, Irenes!", in the example above, are part of a
closely related syntax that uses a period instead of a letter s, and prints
the string out to your terminal instead of returning it.
Whether you're using s" or .", that space after the quote is mandatory,
which may seem very strange if you're more familiar with pretty much any
language that isn't a Forth, but it's a common Forth idiom. Requiring the
space allows the string literal syntax to be tokenized just like any other
space-delimited word. Unlike most languages, there's no special concept of an
operator or punctuation character that can "interrupt" another word or run up
against the start of one. Everything is separated by spaces.
Of course, in modern Forth dialects it's also very common to add a special
lexer feature for that sort of thing. Evocation doesn't do that, because
eventually fancy syntax and grammar will be implemented at a higher layer,
using the experimental parsing formalism that doesn't exist yet. If you want
to see where this feature would be if it were implemented in the simpler way,
you can read the definition of the word named "word", in interpret.e.
If you went to look at that and you're wondering: Yeah, Evocation's lexer
really is that short and simple. Part of why it's able to be that easy is that
words such as s" that introduce special syntaxes do their own lexing for
whatever comes after them.
TODO show how to define words
Evocation has high-level flow-control words: if, unless, if-else, forever,
and while. High-level flow control is a common thing for modern Forth dialects
to add, but every dialect does it a bit differently. Evocation's flow-control
words are postfix operations and work with curly braces, like this:
$ ./evoke
: count 10 0 { 2dup < } { space dup . 1+ } while 2drop newline ;
count
What will it print? :)
Evocation's high-level flow control works only in compiled code; this
example defines and compiles a new word called "count", in order to show it
off. If you try to use the "{ ... } { ... } while" syntax outside of a word
definition, it won't do what you expect.
This is because, unlike some modern Forths, Evocation doesn't have a
general-purpose memory management facility; it uses something called the log,
which makes it easy to allocate things but hard to deallocate them. In order
to loop through a code block, it has to be allocated somewhere. So, the design
takes care not to encourage programming habits that would burn through memory
space.
~~~~~~~~~~~~~~~~~~~~~
~~ Advanced features ~~
~~~~~~~~~~~~~~~~~~~~~
TODO where should this go? should there be an interactive tutorial?
: baz ." baz baz baz!" newline 1 nexit ;
: bar ." bar start" newline baz baz baz ." bar end" newline ;
: foo ." foo start" newline bar ." foo end" newline ;
foo
foo start
bar start
baz baz baz!
baz baz baz!
baz baz baz!
bar end
foo end
: baz ." baz baz baz!" newline 2 nexit ;
: bar ." bar start" newline baz baz baz ." bar end" newline ;
: foo ." foo start" newline bar ." foo end" newline ;
foo
foo start
bar start
baz baz baz!
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
~~ Reading Evocation's source code ~~
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Even if you're only interested in using Evocation, not in modifying it, we
encourage you to at least skim through the source. If you've looked at it,
even a little, it won't be so scary next time. It's heavily commented and
meant for anyone with a little programming knowledge to be able to read, even
if you've never done systems programming before.
If you find something in it confusing, please don't be afraid to ask! It's
likely other people are confused too, and sharing your questions helps improve
the documentation and lets others learn by watching.
The top-level source file whose job is to compile Evocation itself is
evoke.e. It's really short, and worth a quick glance right now. It lists all
the other source files and the order they get loaded in and how they're
processed.
The files that do the work to make Evocation run at all are execution.e and
core.e. It's worth reading through both of them slowly. After you've read
core.e, you'll know a lot of basic words that can be used as commands within
Evocation.
A lot of Evocation is written in Evocation's version of assembly language.
The file amd64.e is the one that implements all the assembly instructions.
Writing a real program in assembly also requires resolving labels, which are
a special syntax that gives names to addresses. The behavior of labels is all
implemented in labels.e.
On the assembly language front, there's also linux.e which contains assembly
words for doing things specific to the Linux operating system, such as reading
input, and there's elf.e which contains words for outputting the special file
headers that let the operating system understand that a file is an executable
program.
In terms of the Forth-y bits, input.e and output.e are concerned with
getting text into and out of the language. The infrastructure to define words
is in dynamic.e, and the syntax for it is in interpret.e. The high-level flow
control words are in flow-control.e. Some of the features of execution.e had
to be separated out into their own file, because of details about how the
compiler works; that stuff is in execution-suport.e.
So, there's all those relatively normal compiler internals in those various
files, which are all fairly self-contained... and then there's the
transformation facility. This is Evocation's most unique archictural decision,
and it's in transform.e. It's well documented, but it's also extremely
conceptually dense. Feel free to give it a skim, that's the only way to build
familiarity with these things, but you should probably have a solid
understanding of the rest of the internals before you place any high
expectations on yourself around understanding the transformation facility.
It's okay, you can benefit from it before you understand it: Transformation
provides the core tricks that make it possible to compile Forth code into
standalone executables. The call to label-transform in evoke.e, and the call
to log-load-transform in execution.e, are the two spots where compilation is
handed off to the transformation facility, and you can pretty much just take
it for granted that it works, until you feel ready.
If you want examples of programs that are smaller than Evocation itself,
quine.e is a tiny program written in proper Evocation that outputs its own
source code; hello.e is a hello-world written in Evocation-assembly, and hex.e
is another small Evocation-assembly program that might make a good example of
how to do slightly more complex things that way. All three of these are
self-contained, consisting of just that one file plus calls to Evocation's
built-in library.
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
~~ Modifying Evocation's Internals ~~
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
There may come a point in your explorations when you wish to make changes to
the compiler. When doing so, please always make sure to build
Evocation-in-Evocation both via Evocation-from-flatassembler, as described in
"Building" above, and with itself, like this:
$ (cat labels.e elf.e transform.e execution.e; echo 's" pyrzqxgl" allocate-string dup 262144 read-to-buffer'; cat core.e linux.e output.e amd64.e execution-support.e log-load.e; echo pyrzqxgl swap 262144 read-to-buffer; cat core.e linux.e output.e amd64.e execution-support.e log-load.e dynamic.e input.e interpret.e flow-control.e linux-dynamic.e ; echo pyrzqxgl; cat evoke.e) | ./evoke > evoke2
$ chmod 755 evoke2
The two versions evoke and evoke2 should be bytewise identical; if they are
not, please fix that. This is an important property which would be very
difficult to get back if we ever lose it, it's easier to maintain it
in-the-moment.
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
~~ Hexing Evocation for Distribution ~~
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
The long-term strategy for Evocation's binary bootstrapping is not yet
ready, but it's described here anyway because this is where the explanation
should eventually go, and it's easier to write about the pieces as they're
created.
The binary bootstrapping strategy rests on something called the
"hex transform", the most complex of the transformations provided as part of
Evocation's transformation facility in transform.e. The hex transform has the
task of transforming an entire compilation process, which would otherwise
produce an executable binary, and instead output a commented hex dump of that
binary which describes its internals and their purpose, byte by byte, in
sufficient detail to allow a human reader to audit their correctnes. It will
do this by passing through comments and call-stack information from the
compilation process to the resulting output.
In order to turn this commented hex dump into a binary, there is a tiny
program called "hex" which handles comments in Evocations ~ syntax, and
converts ASCII hexadecimal to raw binary. This program is in hex.e and is
written in Evocation-assembly. When compiled it is only 480 bytes, which is
small enough to fully audit in its raw, binary form. This is slightly larger
than necessary; many of those bytes are used for error message strings, on the
principle that it's very important that it be easy to distinguish a successful
invocation of "hex" from a failed one.
Although the hex transform is not yet fully the compiled "hex" has proven
quite stable, and the hex transform does work on it. So, a copy of the
compiled "hex" is checked into source control so that it can serve as a root
of trust for future Evocation builds. For ease of auditing, a commented hex
dump version of this binary, produced via the hex transform, is also checked
in, as "hex.hex" (We heard you liked metacircularity, so we put some
metacircularity in your metacircularity so you can be metacircular while
you're metacircular.)
If you need to compile "hex", you can do so as follows:
$ cat labels.e elf.e hex.e | ./evoke > hex
$ chmod 755 hex
To produce the hex-dump version of it, do:
$ (cat labels.e elf.e transform.e; echo 's" xyzzy" allocate-string dup 262144 read-to-buffer'; cat core.e linux.e output.e amd64.e execution-support.e log-load.e dynamic.e input.e interpret.e flow-control.e linux-dynamic.e labels.e elf.e hex.e; echo 'xyzzy s" hex-source" variable 1024 1024 * allocate s" hex-binary" variable 1024 1024 * allocate s" hex-metadata" variable hex-metadata hex-binary dup hex-source 5 roll hex-transform bye ' ) | ./evoke > hex.hex
Although the hex transform doesn't yet work on Forth programs (only programs
written in Evocation-assembly), if you intend to play around with this you may
wish to know how to attempt to run it on things. The latest draft way to do
that is:
$ (cat labels.e elf.e transform.e; echo 's" xyzzy" allocate-string dup 1048576 read-to-buffer'; cat core.e linux.e output.e amd64.e execution-support.e log-load.e dynamic.e input.e interpret.e flow-control.e linux-dynamic.e labels.e elf.e transform.e execution.e; echo 's" pyrzqxgl" allocate-string dup 262144 read-to-buffer '; cat core.e linux.e output.e amd64.e execution-support.e log-load.e; echo pyrzqxgl swap 262144 read-to-buffer; cat core.e linux.e output.e amd64.e execution-support.e log-load.e dynamic.e input.e interpret.e flow-control.e linux-dynamic.e; echo pyrzqxgl; cat evoke.e; echo 'xyzzy s" evoke-source" variable 1024 1024 2 * * allocate s" evoke-binary" variable 1024 1024 4 * * allocate s" evoke-metadata" variable evoke-metadata evoke-binary dup evoke-source 5 roll hex-transform bye ' ) | ./evoke > evoke.hex
It should run to completion, producing output. The output is even correct in
the sense that passing evoke.hex through ./hex will give a binary that's
byte-for-byte identical to evoke, but the output has various problems such as
displaying assembly parameters in the wrong order, having insufficient
explanation of label references and definitions, not showing dictionary entry
headers in any special way, and so on. All these cosmetic issues should be
fixable now, and should likely be the focus of any development efforts.
It also takes several minutes to run, which Irenes believe is because of the
use of linked lists rather than hash tables for the various dictionaries.
Adding a hash table is a task to do after bootstrapping is complete.
Nearly all of this runtime is attributable to the log-load transform; if
you're working on something that doesn't involve the log-load transform, you
may find it useful to temporarily comment out the call to log-load-transform
in execution.e, replacing it with two invocations of drop.
Now get debugging! :)
|