aboutsummaryrefslogtreecommitdiff
path: root/docs/internals.md
blob: b8c81bad9bff1277fb518106c7d92dfbe509589a (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
# Internals

hdass is a straight pipeline: `lex → parse → analyze → emit`. Each stage is one
pair of files under [`src/`](../src).

| Stage | Files | Does |
| --- | --- | --- |
| CLI | `main.c`, `args.c` | parse arguments, pick a target, drive the pipeline |
| Lex | `lexer.c` | source text → a stream of tokens |
| Parse | `parser.c`, `ast.c` | tokens → an AST (`struct Program` of procs and declarations) |
| Analyze | `sema.c` | check the AST: undefined names, entry point, constant/reference rules |
| Emit | `codegen.c` | AST → assembly text for the chosen target |
| Support | `diag.c`, `file.c` | caret diagnostics, file reading |

The AST is mostly architecture-neutral (assignments, control flow, `^` memory,
calls, a raw instruction), so almost all of hdass is shared. The architecture
lives entirely in code generation.

## Two seams in codegen

Code generation is split along the same two axes as a [target](targets.md):

- **`struct Arch`** — instruction selection. One hook, `emit_proc`, turns a
  procedure's statements into that architecture's instructions (its register
  model, mnemonics, stack frames). `x86_arch` and `aarch64_arch` implement it.
- **`struct Backend`** — assembler syntax. Framing hooks (`prologue`, `constant`,
  `data_section`, `string_data`, `float_slot`, `text_section`, `global`,
  `boot_signature`) write the file structure around the instructions.
  `nasm_backend`, `fasm_backend` and `gas_backend` implement it.

`generate(program, out, arch, backend)` orchestrates the two. A public entry
point is just a pairing:

```c
void generate_nasm(struct Program* program, FILE* out)
{
    generate(program, out, &x86_arch, &nasm_backend);
}
```

So `nasm` and `fasm` reuse one x86 instruction selector with different framing,
and `arm64` pairs its own selector with GNU as.

## Adding an assembler backend

To emit a new *syntax* for an existing architecture (say masm for x86-64):

1. Write the framing functions (`masm_prologue`, `masm_constant`, …) and gather
   them into a `static const struct Backend masm_backend`.
2. Add `generate_masm` that pairs `x86_arch` with it.
3. Wire a `-t masm` name in `args.c` and dispatch to it in `main.c`.

Only the framing differs; the instruction bodies come from `x86_arch` unchanged.

## Adding an architecture

To emit a new *instruction set* (the larger job):

1. Write an `emit_proc_<arch>` and the helpers it needs — a register mapping, an
   operand renderer, and lowerings for each statement kind. The AArch64 selector
   is the template: it maps logical `rN → x(N-1)`, renders `#immediate` operands,
   and lowers assignment/arithmetic/branch/call/syscall.
2. Gather it into a `static const struct Arch <arch>_arch`.
3. Pick an assembler `Backend` (GNU as suits most non-x86 targets — reuse
   `gas_backend` or write one), add `generate_<arch>`, and wire a `-t` name.

Unsupported statement kinds should emit a `; TODO` comment instead of incorrect
instructions, the convention the existing selectors already use for gaps.

## Building and checking

Meson drives the build; see [getting started](getting-started.md). The test suite
(`build/tests`) covers the lexer, parser, sema and codegen for every target. If
`cppcheck` is installed, `ninja -C build cppcheck` runs static analysis. The
scripts under [`scripts/`](../scripts) assemble and run every example end to end
in Docker.