From c0f000ea166a24781fcdbf23c080b203cd9b8ed7 Mon Sep 17 00:00:00 2001 From: hachem Date: Fri, 18 Sep 2026 19:47:35 +0200 Subject: chore: update docs --- docs/internals.md | 74 +++++++++++++++++++++++++++---------------------------- 1 file changed, 37 insertions(+), 37 deletions(-) (limited to 'docs/internals.md') diff --git a/docs/internals.md b/docs/internals.md index b8c81ba..f219a8f 100644 --- a/docs/internals.md +++ b/docs/internals.md @@ -1,34 +1,34 @@ -# Internals +# internals -hdass is a straight pipeline: `lex → parse → analyze → emit`. Each stage is one +hdass is a straight pipeline: `lex → parse → analyze → emit`. each stage is one pair of files under [`src/`](../src). -| Stage | Files | Does | +| stage | files | does | | --- | --- | --- | -| CLI | `main.c`, `args.c` | parse arguments, pick a target, drive the pipeline | -| Lex | `lexer.c` | source text → a stream of tokens | -| Parse | `parser.c`, `ast.c` | tokens → an AST (`struct Program` of procs and declarations) | -| Analyze | `sema.c` | check the AST: undefined names, entry point, constant/reference rules | -| Emit | `codegen.c` | AST → assembly text for the chosen target | -| Support | `diag.c`, `file.c` | caret diagnostics, file reading | - -The AST is mostly architecture-neutral (assignments, control flow, `^` memory, -calls, a raw instruction), so almost all of hdass is shared. The architecture +| cli | `main.c`, `args.c` | parse arguments, pick a target, drive the pipeline | +| lex | `lexer.c` | source text → a stream of tokens | +| parse | `parser.c`, `ast.c` | tokens → an ast (`struct Program` of procs and declarations) | +| analyze | `sema.c` | check the ast: undefined names, entry point, constant/reference rules | +| emit | `codegen.c` | ast → assembly text for the chosen target | +| support | `diag.c`, `file.c` | caret diagnostics, file reading | + +the ast is mostly architecture-neutral (assignments, control flow, `^` memory, +calls, a raw instruction), so almost all of hdass is shared. the architecture lives entirely in code generation. -## Two seams in codegen +## two seams in codegen -Code generation is split along the same two axes as a [target](targets.md): +code generation is split along the same two axes as a [target](targets.md): -- **`struct Arch`** — instruction selection. One hook, `emit_proc`, turns a +- **`struct Arch`** — instruction selection. one hook, `emit_proc`, turns a procedure's statements into that architecture's instructions (its register model, mnemonics, stack frames). `x86_arch` and `aarch64_arch` implement it. -- **`struct Backend`** — assembler syntax. Framing hooks (`prologue`, `constant`, +- **`struct Backend`** — assembler syntax. framing hooks (`prologue`, `constant`, `data_section`, `string_data`, `float_slot`, `text_section`, `global`, `boot_signature`) write the file structure around the instructions. `nasm_backend`, `fasm_backend` and `gas_backend` implement it. -`generate(program, out, arch, backend)` orchestrates the two. A public entry +`generate(program, out, arch, backend)` orchestrates the two. a public entry point is just a pairing: ```c @@ -38,39 +38,39 @@ void generate_nasm(struct Program* program, FILE* out) } ``` -So `nasm` and `fasm` reuse one x86 instruction selector with different framing, -and `arm64` pairs its own selector with GNU as. +so `nasm` and `fasm` reuse one x86 instruction selector with different framing, +and `arm64` pairs its own selector with gnu as. -## Adding an assembler backend +## adding an assembler backend -To emit a new *syntax* for an existing architecture (say masm for x86-64): +to emit a new *syntax* for an existing architecture (say masm for x86-64): -1. Write the framing functions (`masm_prologue`, `masm_constant`, …) and gather +1. write the framing functions (`masm_prologue`, `masm_constant`, …) and gather them into a `static const struct Backend masm_backend`. -2. Add `generate_masm` that pairs `x86_arch` with it. -3. Wire a `-t masm` name in `args.c` and dispatch to it in `main.c`. +2. add `generate_masm` that pairs `x86_arch` with it. +3. wire a `-t masm` name in `args.c` and dispatch to it in `main.c`. -Only the framing differs; the instruction bodies come from `x86_arch` unchanged. +only the framing differs; the instruction bodies come from `x86_arch` unchanged. -## Adding an architecture +## adding an architecture -To emit a new *instruction set* (the larger job): +to emit a new *instruction set* (the larger job): -1. Write an `emit_proc_` and the helpers it needs — a register mapping, an - operand renderer, and lowerings for each statement kind. The AArch64 selector +1. write an `emit_proc_` and the helpers it needs — a register mapping, an + operand renderer, and lowerings for each statement kind. the aarch64 selector is the template: it maps logical `rN → x(N-1)`, renders `#immediate` operands, and lowers assignment/arithmetic/branch/call/syscall. -2. Gather it into a `static const struct Arch _arch`. -3. Pick an assembler `Backend` (GNU as suits most non-x86 targets — reuse +2. gather it into a `static const struct Arch _arch`. +3. pick an assembler `Backend` (gnu as suits most non-x86 targets — reuse `gas_backend` or write one), add `generate_`, and wire a `-t` name. -Unsupported statement kinds should emit a `; TODO` comment instead of incorrect +unsupported statement kinds should emit a `; TODO` comment instead of incorrect instructions, the convention the existing selectors already use for gaps. -## Building and checking +## building and checking -Meson drives the build; see [getting started](getting-started.md). The test suite -(`build/tests`) covers the lexer, parser, sema and codegen for every target. If -`cppcheck` is installed, `ninja -C build cppcheck` runs static analysis. The +meson drives the build; see [getting started](getting-started.md). the test suite +(`build/tests`) covers the lexer, parser, sema and codegen for every target. if +`cppcheck` is installed, `ninja -C build cppcheck` runs static analysis. the scripts under [`scripts/`](../scripts) assemble and run every example end to end -in Docker. +in docker. -- cgit v1.3