aboutsummaryrefslogtreecommitdiff
path: root/docs
diff options
context:
space:
mode:
authorhachem <im@hachem.wtf>2026-08-31 03:22:54 +0200
committerhachem <im@hachem.wtf>2026-08-31 03:22:54 +0200
commit81e960be2953a9953fd89c42cc6f71ce6ff54c85 (patch)
treefb60fbe35871cf83ccba75efe13d01b382034240 /docs
parentce9f0d32c33333161cf9a5c2b7446d80fa0f13b5 (diff)
feat: add diagnostics module with source-anchored caret errors
Diffstat (limited to 'docs')
-rw-r--r--docs/language.md163
1 files changed, 163 insertions, 0 deletions
diff --git a/docs/language.md b/docs/language.md
new file mode 100644
index 0000000..a2b4ffa
--- /dev/null
+++ b/docs/language.md
@@ -0,0 +1,163 @@
+# hdass language reference
+
+hdass emits NASM for x86-64; fasm and masm are planned. The compiler output itself isn't tied to an OS, but the examples and toolchain here target Linux (Linux syscall numbers, `nasm -f elf64`, `ld`). Pipeline: `lex → parse → analyze → emit`.
+
+## A first program
+
+```hdass
+[entry: main]
+
+const SYS_WRITE = 1
+const SYS_EXIT = 60
+const STDOUT = 1
+
+data message = "type shi.\n"
+
+proc main
+{
+ rax = SYS_WRITE
+ rdi = STDOUT
+ rsi = message
+ rdx = message.len
+ syscall
+
+ rax = SYS_EXIT
+ rdi = 0
+ syscall
+}
+```
+
+Writes `type shi.` to stdout and exits.
+
+## Comments
+
+```hdass
+rax = 1 // line
+/* block */
+```
+
+## Directives
+
+Top-level `[key: value]`, configuring the whole program.
+
+| Directive | Meaning |
+| --- | --- |
+| `[bits: 64]` / `[bits: 32]` | Target bitness. Default 64. |
+| `[entry: NAME]` | Makes procedure `NAME` the entry point. |
+| `[enable: NAME]` | Turns on an [extension](#extensions). |
+
+Unknown keys, a `bits` value other than 32/64, and unknown extensions are errors.
+
+## Constants and data
+
+`const` names an integer (becomes a NASM `%define`). `data` puts a string in `.data`; the name is its address and `.len` is its length in bytes.
+
+```hdass
+const STDOUT = 1
+data message = "type shi.\n" // message -> address, message.len -> 10
+```
+
+## Procedures
+
+`proc` groups a body. Parameters name registers — `value` below is `rdi`.
+
+```hdass
+proc print_number(value: rdi)
+{
+ rax = value
+}
+```
+
+Each procedure ends with `ret`, except the entry point. `[entry: NAME]` exports `NAME` with `global` and drops its `ret`, so it must end the program itself (an exit syscall). Link with `ld -e NAME`.
+
+## Registers
+
+Written by their architecture names — `rax`–`rdi`, `rbp`, `rsp`, `r8`–`r15` — and their sub-registers (`al`, `ax`, `eax`, `dl`, …), which imply a store's size. The [`logical_registers`](#extensions) extension adds `r1`–`r14`.
+
+## Statements
+
+```hdass
+rax = SYS_WRITE // mov
+rcx -= 1 // += -= *= /= -> add sub imul idiv (/= targets rax)
+rdx = buffer + 31 // address math with + and -
+loop: // label
+goto loop
+if rcx != 0 // == != < <= > >= ; runs the next statement only
+ goto loop
+syscall
+print_number(r12) // call; args go into the callee's parameter registers
+stack buffer[32] // stack buffer; buffer is its base address
+```
+
+## Dereference (`^`)
+
+`^reg` is the memory at the address in `reg` — NASM's `[reg]`. On the left of `=` it stores there. The store width comes from the value operand, so a sized sub-register picks the size:
+
+```hdass
+^rsi = rdx // mov [rsi], rdx (qword)
+^rsi = dl // mov [rsi], dl (byte)
+^rsi = eax // mov [rsi], eax (dword)
+```
+
+A leading size keyword sets the width explicitly. It down-converts a full register to the matching sub-register, and gives an immediate a width NASM would otherwise reject:
+
+```hdass
+^byte rsi = rdx // mov byte [rsi], dl (rdx -> its low byte)
+^word rsi = rax // mov word [rsi], ax
+^dword rsi = r12 // mov dword [rsi], r12d
+^byte rsi = '0' // mov byte [rsi], '0'
+^byte rsi = 10 // mov byte [rsi], 10
+```
+
+`^` is store-only for now; loading with `rax = ^rsi` isn't supported yet.
+
+## Expressions
+
+Assignment values and `if` operands: registers, integers, chars (`'0'`), constants, data names, member access (`data.len`), and `+`/`-`. Multiply and divide come from the `*=` and `/=` compound assignments, not from `*`/`/` inside a value.
+
+## Extensions
+
+### logical_registers
+
+Uniform names for the general-purpose registers, so you don't juggle the irregular `rax`/`rbx`/`rsi`/… spellings. `r1`–`r14` map to:
+
+| `r1` | `r2` | `r3` | `r4` | `r5` | `r6` | `r7` | `r8` | `r9` | `r10` | `r11` | `r12` | `r13` | `r14` |
+| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
+| rax | rbx | rcx | rdx | rsi | rdi | r8 | r9 | r10 | r11 | r12 | r13 | r14 | r15 |
+
+`rsp`/`rbp` and the instruction pointer keep their own names.
+
+```hdass
+r1 = 5 // mov rax, 5
+r4 = r10 // mov rdx, r11
+r6 += r1 // add rdi, rax
+```
+
+A `.8`/`.16`/`.32`/`.64` suffix selects the width, mapping to the sub-register:
+
+```hdass
+r1.8 // al
+r1.16 // ax
+r1.32 // eax
+r1.64 // rax
+^byte rsi = r4 // mov byte [rsi], dl (r4 -> rdx -> dl)
+```
+
+Arch `r8`–`r15` share the `rN` spelling, so with the extension on a bare `r8` is the *logical* register (which is arch `r9`). Reach arch `r8`–`r15` through logical `r7`–`r14`. Architecture names like `rax` and `rsi` still work everywhere.
+
+## Building a program
+
+```bash
+hdass program.hdass -o program.asm
+nasm -f elf64 program.asm -o program.o
+ld -e main program.o -o program
+./program
+```
+
+The [README](../README.md) has a Docker setup with these tools.
+
+## Gotchas
+
+- **Clobbering is yours.** `syscall` trashes `rcx`/`r11`; a callee trashes what it touches. Nothing is saved for you — `examples/fibonacci.hdass` keeps its counter in `r15` for this reason.
+- **Widths must match.** `rax = r1.8` becomes `mov rax, al`, which won't assemble.
+- **The entry procedure has no `ret`** — end it with an exit syscall. \ No newline at end of file