aboutsummaryrefslogtreecommitdiff
diff options
context:
space:
mode:
-rw-r--r--docs/README.md16
-rw-r--r--docs/getting-started.md44
-rw-r--r--docs/internals.md72
-rw-r--r--docs/language.md116
-rw-r--r--docs/targets.md66
5 files changed, 157 insertions, 157 deletions
diff --git a/docs/README.md b/docs/README.md
index fe42c75..c36bd26 100644
--- a/docs/README.md
+++ b/docs/README.md
@@ -2,16 +2,16 @@
hdass is a small superset of assembly: you keep registers, memory and explicit
control flow, and write the repetitive parts (moves, arithmetic, loops, branches)
-more compactly. It transpiles to x86-64 (NASM or fasm) or AArch64 (GNU as).
+more compactly. it transpiles to x86-64 (nasm or fasm) or aarch64 (gnu as).
-- [Getting started](getting-started.md) — build the compiler, then write and run
- a first program on both x86-64 and AArch64.
-- [Language reference](language.md) — the full syntax: declarations, statements,
+- [getting started](getting-started.md) — build the compiler, then write and run
+ a first program on both x86-64 and aarch64.
+- [language reference](language.md) — the full syntax: declarations, statements,
control flow, memory, floating point, raw instructions and extensions.
-- [Targets](targets.md) — the architecture and assembler axes, what each target
- supports, the portable register model, and the per-architecture syscall ABIs.
-- [Internals](internals.md) — the `lex → parse → analyze → emit` pipeline and how
+- [targets](targets.md) — the architecture and assembler axes, what each target
+ supports, the portable register model, and the per-architecture syscall abis.
+- [internals](internals.md) — the `lex → parse → analyze → emit` pipeline and how
to add a new assembler backend or architecture.
-Working programs live in [examples/](../examples/) (x86-64) and
+working programs live in [examples/](../examples/) (x86-64) and
[examples/arm64/](../examples/arm64/).
diff --git a/docs/getting-started.md b/docs/getting-started.md
index 6f90618..6781c0f 100644
--- a/docs/getting-started.md
+++ b/docs/getting-started.md
@@ -1,22 +1,22 @@
-# Getting started
+# getting started
-## Build the compiler
+## build the compiler
-hdass builds with [Meson](https://mesonbuild.com/):
+hdass builds with [meson](https://mesonbuild.com/):
```bash
meson setup build
meson compile -C build
```
-That produces `build/hdass`. It has two options you need: `-o <file>` writes the
+that produces `build/hdass`. it has two options you need: `-o <file>` writes the
output (default stdout), and `-t <target>` picks the assembler and architecture
(`nasm` by default, `fasm`, or `arm64`). `build/hdass --help` lists the rest.
hdass only *transpiles*: it emits assembly text, which you assemble and link
-yourself. The examples target Linux, so you need an x86-64 (and
-for AArch64, an aarch64) Linux toolchain. The bundled
-[Docker image](../README.md#building) has `nasm`, `fasm`, an aarch64
+yourself. the examples target linux, so you need an x86-64 (and
+for aarch64, an aarch64) linux toolchain. the bundled
+[docker image](../README.md#building) has `nasm`, `fasm`, an aarch64
cross-assembler and `qemu`; the commands below run inside it:
```bash
@@ -24,9 +24,9 @@ docker compose up -d
docker compose exec hdass bash
```
-## A first program (x86-64)
+## a first program (x86-64)
-Put this in `hello.hdass`:
+put this in `hello.hdass`:
```hdass
[entry: main]
@@ -51,7 +51,7 @@ proc main
}
```
-Transpile, assemble, link and run:
+transpile, assemble, link and run:
```bash
./build/hdass hello.hdass -o hello.asm
@@ -60,15 +60,15 @@ ld -e main hello.o -o hello
./hello
```
-It prints `Hello, hdass!`. `[entry: main]` exports `main` and drops its `ret`, so
+it prints `Hello, hdass!`. `[entry: main]` exports `main` and drops its `ret`, so
the procedure ends the program itself with the exit syscall; `ld -e main` uses it
-as the start symbol. Swap `-t fasm` and `fasm hello.asm hello.o` to use fasm
+as the start symbol. swap `-t fasm` and `fasm hello.asm hello.o` to use fasm
instead — the machine code is the same.
-## Running on AArch64
+## running on aarch64
-The same source model runs on ARM, but the syscall ABI differs (numbers and
-argument registers), so this program is AArch64-specific. Written with the
+the same source model runs on arm, but the syscall abi differs (numbers and
+argument registers), so this program is aarch64-specific. written with the
[`logical_registers`](targets.md#the-portable-register-model) extension so the
registers read the same on both architectures:
@@ -86,7 +86,7 @@ proc main
}
```
-Build it for AArch64 and run it under qemu:
+build it for aarch64 and run it under qemu:
```bash
./build/hdass -t arm64 exit.hdass -o exit.s
@@ -95,13 +95,13 @@ aarch64-linux-gnu-ld -e main exit.o -o exit
qemu-aarch64 ./exit # exits 42
```
-`r1` maps to `x0` and `r9` to `x8`, `syscall` becomes `svc #0`. See
-[Targets](targets.md) for the register mapping and the ABI tables.
+`r1` maps to `x0` and `r9` to `x8`, `syscall` becomes `svc #0`. see
+[targets](targets.md) for the register mapping and the abi tables.
-## Next
+## next
-- The whole language: [language reference](language.md).
-- What runs where and why programs still carry per-architecture ABI details:
+- the whole language: [language reference](language.md).
+- what runs where and why programs still carry per-architecture abi details:
[targets](targets.md).
-- More programs to read: [examples/](../examples/) and
+- more programs to read: [examples/](../examples/) and
[examples/arm64/](../examples/arm64/), each runnable with the steps above.
diff --git a/docs/internals.md b/docs/internals.md
index b8c81ba..f219a8f 100644
--- a/docs/internals.md
+++ b/docs/internals.md
@@ -1,34 +1,34 @@
-# Internals
+# internals
-hdass is a straight pipeline: `lex → parse → analyze → emit`. Each stage is one
+hdass is a straight pipeline: `lex → parse → analyze → emit`. each stage is one
pair of files under [`src/`](../src).
-| Stage | Files | Does |
+| stage | files | does |
| --- | --- | --- |
-| CLI | `main.c`, `args.c` | parse arguments, pick a target, drive the pipeline |
-| Lex | `lexer.c` | source text → a stream of tokens |
-| Parse | `parser.c`, `ast.c` | tokens → an AST (`struct Program` of procs and declarations) |
-| Analyze | `sema.c` | check the AST: undefined names, entry point, constant/reference rules |
-| Emit | `codegen.c` | AST → assembly text for the chosen target |
-| Support | `diag.c`, `file.c` | caret diagnostics, file reading |
+| cli | `main.c`, `args.c` | parse arguments, pick a target, drive the pipeline |
+| lex | `lexer.c` | source text → a stream of tokens |
+| parse | `parser.c`, `ast.c` | tokens → an ast (`struct Program` of procs and declarations) |
+| analyze | `sema.c` | check the ast: undefined names, entry point, constant/reference rules |
+| emit | `codegen.c` | ast → assembly text for the chosen target |
+| support | `diag.c`, `file.c` | caret diagnostics, file reading |
-The AST is mostly architecture-neutral (assignments, control flow, `^` memory,
-calls, a raw instruction), so almost all of hdass is shared. The architecture
+the ast is mostly architecture-neutral (assignments, control flow, `^` memory,
+calls, a raw instruction), so almost all of hdass is shared. the architecture
lives entirely in code generation.
-## Two seams in codegen
+## two seams in codegen
-Code generation is split along the same two axes as a [target](targets.md):
+code generation is split along the same two axes as a [target](targets.md):
-- **`struct Arch`** — instruction selection. One hook, `emit_proc`, turns a
+- **`struct Arch`** — instruction selection. one hook, `emit_proc`, turns a
procedure's statements into that architecture's instructions (its register
model, mnemonics, stack frames). `x86_arch` and `aarch64_arch` implement it.
-- **`struct Backend`** — assembler syntax. Framing hooks (`prologue`, `constant`,
+- **`struct Backend`** — assembler syntax. framing hooks (`prologue`, `constant`,
`data_section`, `string_data`, `float_slot`, `text_section`, `global`,
`boot_signature`) write the file structure around the instructions.
`nasm_backend`, `fasm_backend` and `gas_backend` implement it.
-`generate(program, out, arch, backend)` orchestrates the two. A public entry
+`generate(program, out, arch, backend)` orchestrates the two. a public entry
point is just a pairing:
```c
@@ -38,39 +38,39 @@ void generate_nasm(struct Program* program, FILE* out)
}
```
-So `nasm` and `fasm` reuse one x86 instruction selector with different framing,
-and `arm64` pairs its own selector with GNU as.
+so `nasm` and `fasm` reuse one x86 instruction selector with different framing,
+and `arm64` pairs its own selector with gnu as.
-## Adding an assembler backend
+## adding an assembler backend
-To emit a new *syntax* for an existing architecture (say masm for x86-64):
+to emit a new *syntax* for an existing architecture (say masm for x86-64):
-1. Write the framing functions (`masm_prologue`, `masm_constant`, …) and gather
+1. write the framing functions (`masm_prologue`, `masm_constant`, …) and gather
them into a `static const struct Backend masm_backend`.
-2. Add `generate_masm` that pairs `x86_arch` with it.
-3. Wire a `-t masm` name in `args.c` and dispatch to it in `main.c`.
+2. add `generate_masm` that pairs `x86_arch` with it.
+3. wire a `-t masm` name in `args.c` and dispatch to it in `main.c`.
-Only the framing differs; the instruction bodies come from `x86_arch` unchanged.
+only the framing differs; the instruction bodies come from `x86_arch` unchanged.
-## Adding an architecture
+## adding an architecture
-To emit a new *instruction set* (the larger job):
+to emit a new *instruction set* (the larger job):
-1. Write an `emit_proc_<arch>` and the helpers it needs — a register mapping, an
- operand renderer, and lowerings for each statement kind. The AArch64 selector
+1. write an `emit_proc_<arch>` and the helpers it needs — a register mapping, an
+ operand renderer, and lowerings for each statement kind. the aarch64 selector
is the template: it maps logical `rN → x(N-1)`, renders `#immediate` operands,
and lowers assignment/arithmetic/branch/call/syscall.
-2. Gather it into a `static const struct Arch <arch>_arch`.
-3. Pick an assembler `Backend` (GNU as suits most non-x86 targets — reuse
+2. gather it into a `static const struct Arch <arch>_arch`.
+3. pick an assembler `Backend` (gnu as suits most non-x86 targets — reuse
`gas_backend` or write one), add `generate_<arch>`, and wire a `-t` name.
-Unsupported statement kinds should emit a `; TODO` comment instead of incorrect
+unsupported statement kinds should emit a `; TODO` comment instead of incorrect
instructions, the convention the existing selectors already use for gaps.
-## Building and checking
+## building and checking
-Meson drives the build; see [getting started](getting-started.md). The test suite
-(`build/tests`) covers the lexer, parser, sema and codegen for every target. If
-`cppcheck` is installed, `ninja -C build cppcheck` runs static analysis. The
+meson drives the build; see [getting started](getting-started.md). the test suite
+(`build/tests`) covers the lexer, parser, sema and codegen for every target. if
+`cppcheck` is installed, `ninja -C build cppcheck` runs static analysis. the
scripts under [`scripts/`](../scripts) assemble and run every example end to end
-in Docker.
+in docker.
diff --git a/docs/language.md b/docs/language.md
index 96f534b..8c10eb4 100644
--- a/docs/language.md
+++ b/docs/language.md
@@ -1,10 +1,10 @@
# hdass language reference
-This is the language reference. hdass transpiles to x86-64 (NASM or fasm, `-t nasm`/`-t fasm`) and AArch64 (`-t arm64`); for which target supports what, the portable [`logical_registers`](#logical_registers) model, and the per-architecture syscall ABIs, see [Targets](targets.md). New here? Start with [Getting started](getting-started.md). Pipeline: `lex → parse → analyze → emit`.
+this is the language reference. hdass transpiles to x86-64 (nasm or fasm, `-t nasm`/`-t fasm`) and aarch64 (`-t arm64`); for which target supports what, the portable [`logical_registers`](#logical_registers) model, and the per-architecture syscall abis, see [targets](targets.md). new here? start with [getting started](getting-started.md). pipeline: `lex → parse → analyze → emit`.
-The examples below use x86-64 register names. The output isn't tied to an OS, but the examples and toolchain here target Linux (ELF, `ld`).
+the examples below use x86-64 register names. the output isn't tied to an os, but the examples and toolchain here target linux (elf, `ld`).
-## A first program
+## a first program
```hdass
[entry: main]
@@ -29,36 +29,36 @@ proc main
}
```
-Writes `type shi.` to stdout and exits.
+writes `type shi.` to stdout and exits.
-## Comments
+## comments
```hdass
rax = 1 // line
/* block */
```
-## Directives
+## directives
-Top-level `[key: value]` (or bare `[key]`), configuring the whole program.
+top-level `[key: value]` (or bare `[key]`), configuring the whole program.
-| Directive | Meaning |
+| directive | meaning |
| --- | --- |
-| `[bits: 64]` / `[bits: 32]` / `[bits: 16]` | Target bitness. Default 64. |
-| `[entry: NAME]` | Makes procedure `NAME` the entry point. |
-| `[enable: NAME]` | Turns on an [extension](#extensions). |
-| `[format: bin]` / `[format: elf]` | Output a flat binary instead of an ELF object. Default elf. |
-| `[org: 0x7C00]` | Set the load address of a flat binary. |
-| `[boot]` | Pad a flat binary to 510 bytes and append the `0xAA55` boot signature. |
+| `[bits: 64]` / `[bits: 32]` / `[bits: 16]` | target bitness. default 64. |
+| `[entry: NAME]` | makes procedure `NAME` the entry point. |
+| `[enable: NAME]` | turns on an [extension](#extensions). |
+| `[format: bin]` / `[format: elf]` | output a flat binary instead of an elf object. default elf. |
+| `[org: 0x7C00]` | set the load address of a flat binary. |
+| `[boot]` | pad a flat binary to 510 bytes and append the `0xAA55` boot signature. |
-The last three build a raw binary instead of a linked ELF, enough for an x86
-boot sector. In `bin` format there are no sections or exported symbols, and code
-comes first so execution starts at the origin. Unknown keys, a bad `bits` value,
+the last three build a raw binary instead of a linked elf, enough for an x86
+boot sector. in `bin` format there are no sections or exported symbols, and code
+comes first so execution starts at the origin. unknown keys, a bad `bits` value,
and unknown extensions are errors.
-## Constants and data
+## constants and data
-`const` names a constant integer expression — integer literals, character literals, other constants, a leading `-`, and `+` `-` `*` `/`. Integers are decimal, `0x` hex, or `0b` binary (these forms work anywhere an integer does). `data` puts a string in `.data`; the name is its address and `.len` is its length in bytes.
+`const` names a constant integer expression — integer literals, character literals, other constants, a leading `-`, and `+` `-` `*` `/`. integers are decimal, `0x` hex, or `0b` binary (these forms work anywhere an integer does). `data` puts a string in `.data`; the name is its address and `.len` is its length in bytes.
```hdass
const STDOUT = 1
@@ -67,9 +67,9 @@ const AREA = 8 * 6 // 48
data message = "type shi.\n" // message -> address, message.len -> 10
```
-## Enums and structs
+## enums and structs
-Both describe compile-time values reached with `Name.member`, which folds to an integer.
+both describe compile-time values reached with `Name.member`, which folds to an integer.
`enum` names a set of constants numbered from 0:
@@ -84,7 +84,7 @@ enum Status
rax = Status.Fail // mov rax, 2
```
-`struct` describes a packed memory layout (no padding). Fields are `name` or `name: size`, where size defaults to `qword`. `Name.field` is the field's byte offset, and `Name.size` is the total size.
+`struct` describes a packed memory layout (no padding). fields are `name` or `name: size`, where size defaults to `qword`. `Name.field` is the field's byte offset, and `Name.size` is the total size.
```hdass
struct Point
@@ -98,11 +98,11 @@ rsi += Point.y // add rsi, 8
rax = Point.size // mov rax, 17
```
-A struct is layout only — it allocates nothing. Pair it with a `stack` buffer sized by `Name.size` and pointer arithmetic (see [examples/records.hdass](../examples/records.hdass)).
+a struct is layout only — it allocates nothing. pair it with a `stack` buffer sized by `Name.size` and pointer arithmetic (see [examples/records.hdass](../examples/records.hdass)).
-## Procedures
+## procedures
-`proc` groups a body. Parameters name registers — `value` below is `rdi`.
+`proc` groups a body. parameters name registers — `value` below is `rdi`.
```hdass
proc print_number(value: rdi)
@@ -111,13 +111,13 @@ proc print_number(value: rdi)
}
```
-Each procedure ends with `ret`, except the entry point. `[entry: NAME]` exports `NAME` with `global` and drops its `ret`, so it must end the program itself (an exit syscall). Link with `ld -e NAME`.
+each procedure ends with `ret`, except the entry point. `[entry: NAME]` exports `NAME` with `global` and drops its `ret`, so it must end the program itself (an exit syscall). link with `ld -e NAME`.
-## Registers
+## registers
-Written by their architecture names — `rax`–`rdi`, `rbp`, `rsp`, `r8`–`r15` — and their sub-registers (`al`, `ax`, `eax`, `dl`, …), which imply a store's size. The [`logical_registers`](#logical_registers) extension adds `r1`–`r14`. The segment registers (`cs ds es fs gs ss`) and control registers (`cr0 cr2 cr3 cr4`) are also recognised, for systems code that sets up segments or switches CPU modes.
+written by their architecture names — `rax`–`rdi`, `rbp`, `rsp`, `r8`–`r15` — and their sub-registers (`al`, `ax`, `eax`, `dl`, …), which imply a store's size. the [`logical_registers`](#logical_registers) extension adds `r1`–`r14`. the segment registers (`cs ds es fs gs ss`) and control registers (`cr0 cr2 cr3 cr4`) are also recognised, for systems code that sets up segments or switches cpu modes.
-## Statements
+## statements
```hdass
rax = SYS_WRITE // mov
@@ -135,9 +135,9 @@ print_number(r12) // call; args go into the callee's parameter registers
stack buf[Point.size] // stack buffer (size is any constant); buf is its base address
```
-## Control flow (`if` / `else` / `while`)
+## control flow (`if` / `else` / `while`)
-`if <expr> <cmp> <expr>` guards either the single next statement or a `{ }` block, and an optional `else` takes its own statement or block. `else if` chains because the `else` body is itself a statement. Comparisons are `==` `!=` `<` `<=` `>` `>=`; a float compare needs an `xmm` register on the left (see [Floating point](#floating-point)).
+`if <expr> <cmp> <expr>` guards either the single next statement or a `{ }` block, and an optional `else` takes its own statement or block. `else if` chains because the `else` body is itself a statement. comparisons are `==` `!=` `<` `<=` `>` `>=`; a float compare needs an `xmm` register on the left (see [floating point](#floating-point)).
```hdass
if rax > rbx
@@ -151,7 +151,7 @@ else
rdi = -1
```
-`while <expr> <cmp> <expr>` runs its statement or `{ }` block for as long as the condition holds, testing it before each pass. It is the same condition as `if`, and desugars to a label, the test, the body, and a jump back — the `loop:`/`goto` you would write by hand. Use `goto` to break out early.
+`while <expr> <cmp> <expr>` runs its statement or `{ }` block for as long as the condition holds, testing it before each pass. it is the same condition as `if`, and desugars to a label, the test, the body, and a jump back — the `loop:`/`goto` you would write by hand. use `goto` to break out early.
```hdass
rbx = 0
@@ -162,22 +162,22 @@ while rcx > 0
}
```
-An optional `.name` right after `while` names the loop's generated labels, so they read as `.name` (top) and `.name_end` (exit) instead of the anonymous `.while_N` — handy for finding a loop in the emitted assembly. Give nested loops distinct names.
+an optional `.name` right after `while` names the loop's generated labels, so they read as `.name` (top) and `.name_end` (exit) instead of the anonymous `.while_N` — handy for finding a loop in the emitted assembly. give nested loops distinct names.
```hdass
while .countdown rcx > 0 // emits `.countdown:` … `jmp .countdown` … `.countdown_end:`
rcx -= 1
```
-A **conditional select** picks one of two register values without a branch: `dst = a if <cond> else b`. It lowers to `csel` on AArch64 (one instruction) and `cmov` on x86 (a default move plus a conditional move, arranged so `dst` may safely alias either source). Both sources must be registers.
+a **conditional select** picks one of two register values without a branch: `dst = a if <cond> else b`. it lowers to `csel` on aarch64 (one instruction) and `cmov` on x86 (a default move plus a conditional move, arranged so `dst` may safely alias either source). both sources must be registers.
```hdass
r3 = r1 if r1 > r2 else r2 // r3 = max(r1, r2), branchless
```
-## Dereference (`^`)
+## dereference (`^`)
-`^reg` is the memory at the address in `reg` — NASM's `[reg]`. On the left of `=` it stores there. The store width comes from the value operand, so a sized sub-register picks the size:
+`^reg` is the memory at the address in `reg` — nasm's `[reg]`. on the left of `=` it stores there. the store width comes from the value operand, so a sized sub-register picks the size:
```hdass
^rsi = rdx // mov [rsi], rdx (qword)
@@ -185,7 +185,7 @@ r3 = r1 if r1 > r2 else r2 // r3 = max(r1, r2), branchless
^rsi = eax // mov [rsi], eax (dword)
```
-A leading size keyword sets the width explicitly. It down-converts a full register to the matching sub-register, and gives an immediate a width NASM would otherwise reject:
+a leading size keyword sets the width explicitly. it down-converts a full register to the matching sub-register, and gives an immediate a width nasm would otherwise reject:
```hdass
^byte rsi = rdx // mov byte [rsi], dl (rdx -> its low byte)
@@ -195,7 +195,7 @@ A leading size keyword sets the width explicitly. It down-converts a full regist
^byte rsi = 10 // mov byte [rsi], 10
```
-`^reg` is also a value — it loads from that address. A size keyword loads a narrower value and zero-extends it into the target:
+`^reg` is also a value — it loads from that address. a size keyword loads a narrower value and zero-extends it into the target:
```hdass
rax = ^rsi // mov rax, [rsi]
@@ -204,7 +204,7 @@ rcx = ^dword rsi // mov ecx, [rsi] (32-bit load zero-extends)
rdx = ^rsi + 4 // load, then add 4
```
-`^signed` before the size sign-extends instead, so a narrower value keeps its sign in the full register. It needs a `byte`, `word`, or `dword` size (a full-width load has nothing to extend):
+`^signed` before the size sign-extends instead, so a narrower value keeps its sign in the full register. it needs a `byte`, `word`, or `dword` size (a full-width load has nothing to extend):
```hdass
rax = ^signed byte rsi // movsx rax, byte [rsi]
@@ -212,9 +212,9 @@ rbx = ^signed word rsi // movsx rbx, word [rsi]
rcx = ^signed dword rsi // movsxd rcx, dword [rsi]
```
-## Raw instructions
+## raw instructions
-Any statement that isn't an assignment, label, call or keyword is emitted as a bare instruction: a mnemonic and its operands, written in hdass's own operand syntax. This is the escape hatch for everything outside the assignment and control-flow model: `int`, `hlt`, `cli`/`sti`, `lgdt`, port I/O, the mode switch. Operands are the usual registers, immediates, constants and `^memory`, and share the mnemonic's line.
+any statement that isn't an assignment, label, call or keyword is emitted as a bare instruction: a mnemonic and its operands, written in hdass's own operand syntax. this is the escape hatch for everything outside the assignment and control-flow model: `int`, `hlt`, `cli`/`sti`, `lgdt`, port i/o, the mode switch. operands are the usual registers, immediates, constants and `^memory`, and share the mnemonic's line.
```hdass
cli
@@ -225,19 +225,19 @@ cr0 = eax // an ordinary move; segment/control regs work with `=` too
hlt
```
-Mnemonics pass straight through, so a typo is reported by the assembler. Operands render through the same path as everywhere else, so an instruction is as portable as the rest of the language, though the mnemonics themselves are architecture-specific.
+mnemonics pass straight through, so a typo is reported by the assembler. operands render through the same path as everywhere else, so an instruction is as portable as the rest of the language, though the mnemonics themselves are architecture-specific.
-## Expressions
+## expressions
-Assignment values and `if` operands: registers, integers, chars (`'0'`), constants, data names, member access (`data.len`), and `+` `-` `*` `/` `%`. Operators are left-associative and each right-hand operand must be a single term, so `a * b + c` works but `a + b * c` (a nested right operand) doesn't yet.
+assignment values and `if` operands: registers, integers, chars (`'0'`), constants, data names, member access (`data.len`), and `+` `-` `*` `/` `%`. operators are left-associative and each right-hand operand must be a single term, so `a * b + c` works but `a + b * c` (a nested right operand) doesn't yet.
-A leading `-` negates a term (`rax = -5`, `rbx = rax + -3`, `const OFFSET = -8`). It only applies to values that fold to a constant, so it emits a negative immediate; negating a register (`-rbx`) is not supported.
+a leading `-` negates a term (`rax = -5`, `rbx = rax + -3`, `const OFFSET = -8`). it only applies to values that fold to a constant, so it emits a negative immediate; negating a register (`-rbx`) is not supported.
-## Extensions
+## extensions
### logical_registers
-Uniform names for the general-purpose registers, so you don't juggle the irregular `rax`/`rbx`/`rsi`/… spellings. `r1`–`r14` map to:
+uniform names for the general-purpose registers, so you don't juggle the irregular `rax`/`rbx`/`rsi`/… spellings. `r1`–`r14` map to:
| `r1` | `r2` | `r3` | `r4` | `r5` | `r6` | `r7` | `r8` | `r9` | `r10` | `r11` | `r12` | `r13` | `r14` |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
@@ -251,7 +251,7 @@ r4 = r10 // mov rdx, r11
r6 += r1 // add rdi, rax
```
-A `.8`/`.16`/`.32`/`.64` suffix selects the width, mapping to the sub-register:
+a `.8`/`.16`/`.32`/`.64` suffix selects the width, mapping to the sub-register:
```hdass
r1.8 // al
@@ -261,18 +261,18 @@ r1.64 // rax
^byte rsi = r4 // mov byte [rsi], dl (r4 -> rdx -> dl)
```
-Arch `r8`–`r15` share the `rN` spelling, so with the extension on a bare `r8` is the *logical* register (which is arch `r9`). Reach arch `r8`–`r15` through logical `r7`–`r14`. Architecture names like `rax` and `rsi` still work everywhere.
+arch `r8`–`r15` share the `rN` spelling, so with the extension on a bare `r8` is the *logical* register (which is arch `r9`). reach arch `r8`–`r15` through logical `r7`–`r14`. architecture names like `rax` and `rsi` still work everywhere.
-## Floating point
+## floating point
-Floating-point values live in the SSE registers `xmm0`–`xmm15` (double precision). Float literals like `3.14` are placed in `.data` and loaded for you.
+floating-point values live in the sse registers `xmm0`–`xmm15` (double precision). float literals like `3.14` are placed in `.data` and loaded for you.
```hdass
xmm0 = 3.5 // movsd from a .data slot
xmm0 *= xmm1 // += -= *= /= -> addsd subsd mulsd divsd
```
-An `=` between a float register and a general-purpose register converts:
+an `=` between a float register and a general-purpose register converts:
```hdass
xmm0 = rax // int -> float (cvtsi2sd)
@@ -293,9 +293,9 @@ if xmm0 > 4.0
goto escaped
```
-See [examples/mandelbrot.hdass](../examples/mandelbrot.hdass) for a float program. Not yet supported: mixing floats and ints in one expression, and printing floats.
+see [examples/mandelbrot.hdass](../examples/mandelbrot.hdass) for a float program. not yet supported: mixing floats and ints in one expression, and printing floats.
-## Building a program
+## building a program
```bash
hdass program.hdass -o program.asm # nasm (default)
@@ -304,7 +304,7 @@ ld -e main program.o -o program
./program
```
-Or target fasm with `-t fasm`, which assembles in one step:
+or target fasm with `-t fasm`, which assembles in one step:
```bash
hdass -t fasm program.hdass -o program.asm
@@ -312,7 +312,7 @@ fasm program.asm program.o
ld -e main program.o -o program
```
-For `-t arm64` and the full toolchain (including the AArch64 cross-assembler and qemu), see [Getting started](getting-started.md); the [README](../README.md) has the Docker setup.
+for `-t arm64` and the full toolchain (including the aarch64 cross-assembler and qemu), see [getting started](getting-started.md); the [readme](../README.md) has the docker setup.
-## Some stinkies
-Clobbering is your responsibility: `syscall` trashes `rcx` and `r11`, while a callee can trash any registers it touches, so nothing is saved automatically. `examples/fibonacci.hdass`, for example, keeps its counter in `r15` for this reason. Register widths must also match, meaning something like `rax = r1.8` would become `mov rax, al`, which will not assemble. Division clobbers extra registers: `/` `%` and their `=` forms use `idiv` through `rax:rdx`, so both are overwritten regardless of the destination. The divisor can be anything — a register, a constant, or an immediate — but an immediate or an `rax`/`rdx` divisor is first copied into `r11`, so those also clobber `r11`. Labels and procedures become plain assembler symbols, so avoid names the target assembler reserves: `loop`, for instance, is an instruction mnemonic that fasm rejects as a label (nasm allows it). Finally, the entry procedure has no `ret`; it should end with an exit syscall.
+## some stinkies
+clobbering is your responsibility: `syscall` trashes `rcx` and `r11`, while a callee can trash any registers it touches, so nothing is saved automatically. `examples/fibonacci.hdass`, for example, keeps its counter in `r15` for this reason. register widths must also match, meaning something like `rax = r1.8` would become `mov rax, al`, which will not assemble. division clobbers extra registers: `/` `%` and their `=` forms use `idiv` through `rax:rdx`, so both are overwritten regardless of the destination. the divisor can be anything — a register, a constant, or an immediate — but an immediate or an `rax`/`rdx` divisor is first copied into `r11`, so those also clobber `r11`. labels and procedures become plain assembler symbols, so avoid names the target assembler reserves: `loop`, for instance, is an instruction mnemonic that fasm rejects as a label (nasm allows it). finally, the entry procedure has no `ret`; it should end with an exit syscall.
diff --git a/docs/targets.md b/docs/targets.md
index c54636b..70259ea 100644
--- a/docs/targets.md
+++ b/docs/targets.md
@@ -1,66 +1,66 @@
-# Targets
+# targets
hdass separates two things that assemblers usually tangle together:
- the **architecture** — which instructions exist and how registers work;
- the **assembler syntax** — how those instructions are written to a file.
-A target is a pairing of the two, chosen with `-t`:
+a target is a pairing of the two, chosen with `-t`:
| `-t` | architecture | assembler | notes |
| --- | --- | --- | --- |
-| `nasm` (default) | x86-64 | NASM | Intel syntax, `nasm -f elf64` |
-| `fasm` | x86-64 | fasm | Intel syntax, `fasm` (one step) |
-| `arm64` | AArch64 | GNU as | `aarch64-linux-gnu-as` |
+| `nasm` (default) | x86-64 | nasm | intel syntax, `nasm -f elf64` |
+| `fasm` | x86-64 | fasm | intel syntax, `fasm` (one step) |
+| `arm64` | aarch64 | gnu as | `aarch64-linux-gnu-as` |
-For x86-64, `nasm` and `fasm` emit the **same instruction bodies** and differ only
+for x86-64, `nasm` and `fasm` emit the **same instruction bodies** and differ only
in framing (file header, sections, constant and data syntax). `arm64` is a
separate instruction selector: different registers, three-operand arithmetic,
`ldr`/`str` memory, `cmp`+`b.cond` branches and `svc #0` syscalls.
`masm` (x86-64) and a 32-bit `arm` target are planned.
-## What each architecture supports
+## what each architecture supports
-The language is the same; not every construct lowers on every architecture yet.
+the language is the same; not every construct lowers on every architecture yet.
-| Feature | x86-64 | AArch64 |
+| feature | x86-64 | aarch64 |
| --- | --- | --- |
-| Moves, arithmetic (`+ - * /`), compound assignment | ✅ | ✅ |
+| moves, arithmetic (`+ - * /`), compound assignment | ✅ | ✅ |
| `if` / `else` / `while`, `goto`, labels | ✅ | ✅ |
-| Conditional select (`a if c else b`) | ✅ `cmov` | ✅ `csel` |
-| Calls, `syscall` | ✅ | ✅ |
-| Memory load/store (`^`), sized and signed | ✅ | partial (`ldr`/`str`) |
-| Raw instruction statement | ✅ | ✅ |
-| Modulo (`%`), division remainder | ✅ | ❌ not yet |
-| Floating point (`xmm`) | ✅ | ❌ not yet |
+| conditional select (`a if c else b`) | ✅ `cmov` | ✅ `csel` |
+| calls, `syscall` | ✅ | ✅ |
+| memory load/store (`^`), sized and signed | ✅ | partial (`ldr`/`str`) |
+| raw instruction statement | ✅ | ✅ |
+| modulo (`%`), division remainder | ✅ | ❌ not yet |
+| floating point (`xmm`) | ✅ | ❌ not yet |
| `stack` buffers | ✅ | ❌ not yet |
-| Bare-metal directives (`format`, `org`, `boot`, `bits 16`) | ✅ | — (x86/BIOS concept) |
+| bare-metal directives (`format`, `org`, `boot`, `bits 16`) | ✅ | — (x86/bios concept) |
-Unsupported constructs emit a `; TODO` comment instead of incorrect instructions.
+unsupported constructs emit a `; TODO` comment instead of incorrect instructions.
-## The portable register model
+## the portable register model
-Architecture-native register names (`rax` on x86-64, `x0` on AArch64) lock a
-program to one architecture. To write for both, enable
+architecture-native register names (`rax` on x86-64, `x0` on aarch64) lock a
+program to one architecture. to write for both, enable
[`logical_registers`](language.md#logical_registers): `r1`–`r14` are the
general-purpose registers, mapped per target.
| logical | `r1` | `r2` | `r3` | `r4` | `r5` | `r6` | `r7` | `r8` | `r9` | `r10` | … |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| x86-64 | rax | rbx | rcx | rdx | rsi | rdi | r8 | r9 | r10 | r11 | … |
-| AArch64 | x0 | x1 | x2 | x3 | x4 | x5 | x6 | x7 | x8 | x9 | … |
+| aarch64 | x0 | x1 | x2 | x3 | x4 | x5 | x6 | x7 | x8 | x9 | … |
-AArch64 is simply `rN → x(N-1)`. The raw [instruction statement](language.md#raw-instructions)
+aarch64 is simply `rN → x(N-1)`. the raw [instruction statement](language.md#raw-instructions)
is architecture-locked too: its mnemonics are whatever you write.
-## Syscall ABIs differ
+## syscall abis differ
-Even with logical registers, a *syscall* is not portable: Linux uses different
-call numbers, argument registers and trap instructions per architecture. So a
-program still carries architecture-specific ABI constants.
+even with logical registers, a *syscall* is not portable: linux uses different
+call numbers, argument registers and trap instructions per architecture. so a
+program still carries architecture-specific abi constants.
-| | x86-64 | AArch64 |
+| | x86-64 | aarch64 |
| --- | --- | --- |
| syscall number in | `rax` (logical `r1`) | `x8` (logical `r9`) |
| arguments in | `rdi rsi rdx r10 r8 r9` | `x0 x1 x2 x3 x4 x5` |
@@ -68,10 +68,10 @@ program still carries architecture-specific ABI constants.
| `exit` number | `60` | `93` |
| `write` number | `1` | `64` |
-C works the same way: portable source, per-platform syscalls.
+c works the same way: portable source, per-platform syscalls.
-## OS independence
+## os independence
-The emitted instructions aren't tied to an OS; only the syscall numbers and the
-`[entry]`/link convention are. The examples and toolchain here target Linux (ELF,
-`ld`, and `qemu-aarch64` for ARM); see [getting started](getting-started.md).
+the emitted instructions aren't tied to an os; only the syscall numbers and the
+`[entry]`/link convention are. the examples and toolchain here target linux (elf,
+`ld`, and `qemu-aarch64` for arm); see [getting started](getting-started.md).