aboutsummaryrefslogtreecommitdiff
path: root/docs/internals.md
diff options
context:
space:
mode:
authorhachem <im@hachem.wtf>2026-09-13 00:37:24 +0200
committerhachem <im@hachem.wtf>2026-09-13 00:37:24 +0200
commit4e1edee7383af381d51b11fce45b6a15b62f10e7 (patch)
tree790c0af3701a732d191891c489b2644dc6d45003 /docs/internals.md
parentd50af2e474281a1e88d191f783c2792824698974 (diff)
docs: split docs
Diffstat (limited to 'docs/internals.md')
-rw-r--r--docs/internals.md76
1 files changed, 76 insertions, 0 deletions
diff --git a/docs/internals.md b/docs/internals.md
new file mode 100644
index 0000000..b8c81ba
--- /dev/null
+++ b/docs/internals.md
@@ -0,0 +1,76 @@
+# Internals
+
+hdass is a straight pipeline: `lex → parse → analyze → emit`. Each stage is one
+pair of files under [`src/`](../src).
+
+| Stage | Files | Does |
+| --- | --- | --- |
+| CLI | `main.c`, `args.c` | parse arguments, pick a target, drive the pipeline |
+| Lex | `lexer.c` | source text → a stream of tokens |
+| Parse | `parser.c`, `ast.c` | tokens → an AST (`struct Program` of procs and declarations) |
+| Analyze | `sema.c` | check the AST: undefined names, entry point, constant/reference rules |
+| Emit | `codegen.c` | AST → assembly text for the chosen target |
+| Support | `diag.c`, `file.c` | caret diagnostics, file reading |
+
+The AST is mostly architecture-neutral (assignments, control flow, `^` memory,
+calls, a raw instruction), so almost all of hdass is shared. The architecture
+lives entirely in code generation.
+
+## Two seams in codegen
+
+Code generation is split along the same two axes as a [target](targets.md):
+
+- **`struct Arch`** — instruction selection. One hook, `emit_proc`, turns a
+ procedure's statements into that architecture's instructions (its register
+ model, mnemonics, stack frames). `x86_arch` and `aarch64_arch` implement it.
+- **`struct Backend`** — assembler syntax. Framing hooks (`prologue`, `constant`,
+ `data_section`, `string_data`, `float_slot`, `text_section`, `global`,
+ `boot_signature`) write the file structure around the instructions.
+ `nasm_backend`, `fasm_backend` and `gas_backend` implement it.
+
+`generate(program, out, arch, backend)` orchestrates the two. A public entry
+point is just a pairing:
+
+```c
+void generate_nasm(struct Program* program, FILE* out)
+{
+ generate(program, out, &x86_arch, &nasm_backend);
+}
+```
+
+So `nasm` and `fasm` reuse one x86 instruction selector with different framing,
+and `arm64` pairs its own selector with GNU as.
+
+## Adding an assembler backend
+
+To emit a new *syntax* for an existing architecture (say masm for x86-64):
+
+1. Write the framing functions (`masm_prologue`, `masm_constant`, …) and gather
+ them into a `static const struct Backend masm_backend`.
+2. Add `generate_masm` that pairs `x86_arch` with it.
+3. Wire a `-t masm` name in `args.c` and dispatch to it in `main.c`.
+
+Only the framing differs; the instruction bodies come from `x86_arch` unchanged.
+
+## Adding an architecture
+
+To emit a new *instruction set* (the larger job):
+
+1. Write an `emit_proc_<arch>` and the helpers it needs — a register mapping, an
+ operand renderer, and lowerings for each statement kind. The AArch64 selector
+ is the template: it maps logical `rN → x(N-1)`, renders `#immediate` operands,
+ and lowers assignment/arithmetic/branch/call/syscall.
+2. Gather it into a `static const struct Arch <arch>_arch`.
+3. Pick an assembler `Backend` (GNU as suits most non-x86 targets — reuse
+ `gas_backend` or write one), add `generate_<arch>`, and wire a `-t` name.
+
+Unsupported statement kinds should emit a `; TODO` comment instead of incorrect
+instructions, the convention the existing selectors already use for gaps.
+
+## Building and checking
+
+Meson drives the build; see [getting started](getting-started.md). The test suite
+(`build/tests`) covers the lexer, parser, sema and codegen for every target. If
+`cppcheck` is installed, `ninja -C build cppcheck` runs static analysis. The
+scripts under [`scripts/`](../scripts) assemble and run every example end to end
+in Docker.