Chapter 11
Calling C
A Rill function may have any name, C's own included: fn read, fn bind, fn connect are the program's, since no Rill function is a symbol anything outside the object can see — only export fn and main are — and the runtime's calls to C reach C. The one name a program cannot give a function is one it also declares extern, which the checker refuses.
An extern declaration names a C symbol and its types. The optional string is a library to link with; leave it out for anything already in the C runtime. A string starting with - is passed to the linker as written, and framework:Name links a macOS framework:
extern "-L/opt/homebrew/lib" fn glfwInit() -> I32 # where to search extern "glfw" fn glfwPollEvents() # -lglfw extern "framework:OpenGL" fn glClear(mask: U32) # -framework OpenGL
extern "m" fn sqrt(x: Float) -> Float extern fn getpid() -> Int extern fn strlen(s: Ptr) -> Int
Only shapes the C ABI carries may appear: Int, Float, Bool, Ptr and Unit. Ptr is an opaque C pointer that reference counting leaves alone — whatever produced it owns it.
Int and Float mean C's 64-bit forms, and most C libraries are not written in those. A foreign signature can name the exact width the symbol was compiled with — I8, U8, I16, U16, I32, U32, I64, U64, F32, F64 — and the call narrows arguments on the way in and widens results on the way back. The Rill side is unaffected: an I32 parameter still takes an ordinary Int, and an F32 result still comes back as a Float.
extern "m" fn sqrtf(x: F32) -> F32 # C `float`, not `double` extern fn abs(n: I32) -> I32 # C `int`, not a 64-bit one
One width has no number to name it: C's size_t is as wide as a pointer, which is 64 bits everywhere Rill links natively and 32 on wasm32. USize and ISize are that width, filled in by whatever is being built for, so one declaration is right on both:
extern fn strlen(s: Ptr) -> USize extern fn qsort(base: Ptr, n: USize, size: USize, cmp: Fn(Ptr, Ptr) -> I32) -> Unit
They are for signatures, not storage: Buf refuses them, because a buffer's layout is the program's to decide and this width is the target's.
Getting this wrong is not a compile error but a silently wrong answer, and the float case is the sharp one: a float parameter handed a double reads the wrong half of the register. Width names are only spellable inside an extern.
Raw buffers cover the other half of a C API — the array a library wants to read or fill. They are plain C memory outside the reference counter, so they are released with free_ptr, and the index is in elements, not bytes.
alloc_bytes(n) -> Ptr |
n zeroed bytes, or null if n <= 0 |
ptr_offset(p, bytes) -> Ptr |
byte-level addressing, when a layout needs it |
poke_i8/i16/i32/i64(p, i, v) |
store an Int, narrowed to that width |
poke_f32/f64(p, i, v) |
store a Float |
poke_ptr(p, i, q) / peek_ptr(p, i) -> Ptr |
a pointer slot: char** in, a handle written by an out-parameter back |
peek_i8/i16/i32/i64(p, i) -> Int |
read back, sign-extended |
peek_u8/u16/u32(p, i) -> Int |
read back, zero-extended |
peek_f32/f64(p, i) -> Float |
read back a float |
v = alloc_bytes(3 * 4) # three 32-bit floats poke_f32(v, 0, 1.5) poke_f32(v, 1, 2.5) draw(v, 3) # an extern taking (Ptr, I32) free_ptr(v)
Nothing checks the bounds — a buffer is a length you keep track of yourself, exactly as in C.
Buffers
Raw pointers are how C is reached, but inside Rill the same memory is better had as a Buf(w): a counted block of elements of one C width.
v = buf_f32(3 * 5) # fifteen 32-bit floats, zeroed v[0] = 1.5 println(v[0]) println(buf_len(v)) draw(v, 3) # an extern taking (Ptr, I32)
#v is how many elements it has, v[i] and v[i] = x are one of them; buf_get(v, i) and buf_set(v, i, x) are the same two operations spelled out, and either may be written. Indexing binds tighter than every operator, so v[i + 1] * 2 reads the way it looks, and v[i] = x may only appear as a statement — it is the one assignment the language has, and only to a buffer element.
buf_i8 buf_u8 buf_i16 buf_u16 buf_i32 buf_u32 buf_int buf_f32 buf_float buf_ptr |
one constructor per element width; the argument is a length in elements |
b[i] / b[i] = v — or buf_get(b, i) / buf_set(b, i, v) |
one element, bounds-checked |
#b (or buf_len(b)) |
how many elements |
buf_copy(dst, to, src, from, n) |
n elements of src from from into dst at to, overlapping or not, both bounds-checked |
buf_cmp(a, at, b, bt, n) |
n elements of each as bytes: -1, 0 or 1, both bounds-checked |
buf_copy and buf_cmp are memmove and memcmp with the question asked first: a length that came from somewhere it should not have — a header on a page that is not a header — stops at the check, with a message, rather than in the memory past the buffer. They cost the check and nothing else, and the check is nothing beside the copy; a program that reaches for ptr_offset and an extern memmove to move bytes between buffers has no reason to.
Where a Ptr is wanted — a foreign parameter, a peek/poke, a Rill parameter annotated Ptr — a buffer written there is its first element's address (which is also what the value is underneath: the count and the length live in the sixteen bytes before the elements). This is C's own rule for an array, and it is the only place in the language where a value is not the type it was written as. Nothing about the buffer changes: it stays counted, stays the caller's, and a buffer built in the argument itself is released after the call rather than before, so the address is good for as long as the call is.
b = buf_u8(8) strlen(b) # extern fn strlen(s: Ptr) -> USize from_cstr(b) # a builtin taking a Ptr peek_u8(b, 1) # so are peek and poke
Three things follow from the width living in the type. Reading gives back the Rill type that width stands for, so Buf(U8) yields 200 where Buf(I8) yields -56 for the same byte. An index outside the buffer ends the program with the index and the length, rather than reading whatever was next in memory. A for loop that indexes a buffer by its own variable — v[i] for i in lo..hi — is tested once, before the loop, for lo >= 0 and hi <= #v, and not again at any element; an out-of-range loop stops before it starts, naming the index it would have reached. And a buffer is reference counted like a string — it is released when the last reference goes, so there is no free to forget.
The width is never inferred, since it is not a type variable: a function that takes a buffer says which one, fn fill(v: Buf(F32), i) = .... This is also why a buffer should be held by a Rill binding rather than stashed in a raw Ptr slot, which would lose the count.
examples/buffers.rill shows all of it, and examples/gl/scene.rill fills one with the vertex data a GPU is handed. For a buffer read as a column — scanned in ranges and reduced — lib/num/col.rill has the aggregates written with their accumulators in lanes: import "num/col" as c, then c.sum(v, lo, hi), c.mean, c.variance, c.min_of, c.max_of, c.dot. For arrays with a shape — matrices, vectors, the things NumPy does — lib/num/grid.rill lays a shape and strides over one such buffer: import "num/grid" as g, then g.matmul, g.sum_axis, g.transpose (a view, no copy) and the rest; benchmarks/numpy/README.md lists what it has and how it measures.
Callbacks
Fn(...) in a foreign signature is a C function pointer, and what may be passed for one is a top-level function named directly:
extern fn signal(sig: I32, handler: Fn(I32)) -> Ptr fn on_signal(n) = println("signal " + int_to_str(n)) fn main() = signal(30, on_signal)
A callback that answers writes its result after an arrow, and the same narrowing runs in reverse on the way back out:
extern fn qsort(base: Ptr, n: U64, size: U64, cmp: Fn(Ptr, Ptr) -> I32) -> Unit fn by_size(a, b) = peek_i32(b, 0) - peek_i32(a, 0)
Without the arrow the callback hands nothing back, which is the shape most notification callbacks want. The result may be any type C can carry — a number, a Bool or a Ptr — but not a Str or anything else on the heap, which C would discard with nothing left to release it, and not another Fn.
The compiler wraps the function in a C-ABI trampoline that widens C's arguments to Rill's, so the callback is written in ordinary types.
A callback may also travel as an ordinary parameter, so a helper can wrap the registration rather than every call site repeating it:
fn sort_by(b: Buf(I32), cmp) = qsort(b, buf_len(b), 4, cmp) b fn main() = sort_by(nums, ascending)
What crosses is still a bare code pointer, so the callback has to be settled while compiling: the caller names it, and sort_by is emitted once per comparator it is called with, each copy with its own trampoline built in. A helper like this cannot itself be used as a value, since it exists in as many copies as it has callers. A callback must not block on a channel, since it runs on C's call stack rather than on its own.
A lambda or closure has an environment, and a bare code pointer has nowhere to keep it. Most APIs that take a callback also take a pointer they hand back to it unread — qsort_r's arg, pthread_create's, a library's "user data" — and that is where it goes. The signature says so with Env: the parameter the pointer goes out by, and the slot of the callback's Fn(...) it comes back into.
extern fn rill_sum_by(n: I64, f: Fn(I64, Env) -> I64, arg: Env) -> I64 fn main() = k = 10 println(int_to_str(rill_sum_by(4, \i -> i * k))) # 60
The Rill call leaves the Env out — it is the compiler's to fill — and the callback is written in the types that remain, (Int) -> Int here. What goes out is the address of a box holding the closure, code and environment both; what C calls is a trampoline made for the C signature, which unpacks the box and calls the closure the way every closure is called. The box, and the frame's hold on the closure, last the call: this is for a callback C uses during the call, as a comparator or a visitor is. A callback C keeps after the call returns — a handler registered once and called later — is still a named function, since nothing here can know how long C keeps it. A named function or a value chosen at run time may be passed for an Env callback too. An extern has one Env parameter and one callback carrying it, or neither.
examples/callbacks.rill has C calling back during main, after it has returned, and into a comparator that reaches C through such a helper. examples/ffi_buffers.rill exercises all of it, and examples/gl/ drives OpenGL through nothing else: a window, shaders, vertex data and a spinning triangle, with no C written anywhere.
A Rill Str is a counted heap object rather than a char*, so strings cross the boundary explicitly:
to_cstr(s) -> Ptr |
NUL-terminated copy; free it with free_ptr |
from_cstr(p) -> Str |
copies a C string back into a Rill one |
str_from_ptr(p, n) -> Str |
the same by length rather than by terminator, for bytes |
free_ptr(p) |
frees what to_cstr (or C malloc) returned |
null_ptr() / is_null(p) |
for APIs that use null |
from_cstr stops at the first zero, which is right for a C string and wrong for everything a library hands back by filling a buffer and returning a count. str_from_ptr(p, n) takes exactly n bytes, zeros included, which is what a read, a decompression or an SSL_read gives:
b = buf_u8(16384) n = SSL_read(ssl, b, 16384) if n > 0 then str_from_ptr(b, n) else ""
fn env(name) = p = to_cstr(name) v = getenv(p) free_ptr(p) if is_null(v) then "<unset>" else from_cstr(v)
Two things to keep in mind: a wrong signature is a crash, not a type error — the compiler believes what you declare; and a C call that blocks holds the worker thread running it, since the scheduler cannot preempt foreign code — every other strand on that worker waits with it. A call that may wait — a read, a sleep, a name lookup — is declared extern blocking fn, and is then made on a helper thread while the strand is parked, the worker going on with the others; the answer comes back as if the call had been made here. What it costs is a hand-over each way, a few microseconds, which is why it is said per declaration and not assumed. A blocking call cannot take a callback, since the callback would run on the helper thread, where Rill code cannot; a program without strands makes the call itself.
extern blocking fn usleep(us: U32) -> I32 # the other strands run meanwhile
Being called: export, WebAssembly and shared libraries
extern is Rill calling out. export is the other direction — a function the embedding host may call in:
export fn frame(t: Float) = ...
It means something where there is a host: a wasm module (--target wasm32-wasip1), or a shared library (--shared) that a C or Python program loads. A native executable has neither and refuses the word. Either way the module links as a reactor: it has no entry point of its own, it stays loaded, and the host drives it. main still runs — as rill_init, which the host calls once before anything else — and each export fn is a symbol under its own name.
On wasm the boundary takes what wasm can pass: parameters must be Int, Float, Bool or Ptr, the return those or Unit, and the function must be concrete — an export is one entry point, not a generic family. Data wider than a scalar crosses through a Buf: build one, hand it to an import — where a Ptr is wanted a buffer is its address — and the host reads it straight out of the module's linear memory.
A shared library's boundary is wider: a counted value — a data value, a buffer, a string, a map — crosses as a handle. One the host passes in is lent for the call and is the host's still when the call returns; one an export returns is the host's to keep, and for every such type the library also exports rill_drop_<Type>(handle) — rill_drop_g_Grid for a Grid imported as g, the module's dot made an underscore — for the host to let it go by. A buffer's handle is the address of its elements, so the host may read them straight through it while it holds the grid. bindings/python/gridlib.rill is a library of this kind over lib/num/grid.rill, and bindings/python/grid the ctypes package that makes it import grid as g in Python, NumPy's notation on Rill's loops. Strands need the scheduler a program's main runs and a library's host does not, so a library runs on one worker.
On wasm an extern fn that no library satisfies becomes an import — a hole the host fills at instantiation. That is the whole FFI story in a browser: declare extern fn gpu_draw(v: Ptr, n: Int), supply env.gpu_draw from JavaScript. examples/webgpu/ renders through WebGPU exactly this way — the host owns the asynchronous device setup that a wasm module cannot await, Rill owns the frame.
What a wasm build will not do is link a library by name: a module imports from its host rather than loading a shared object, so extern "glfw" is refused where it would otherwise fail obscurely at run time. "m" and "c" are the exception, and not as a courtesy — wasi-libc is a single archive with the maths already in it, so a program written against -lm on a Mac needs no second spelling for the browser. A C name that wasi genuinely lacks — getpid, signal, anything about processes or users — is a link error naming it, and a declaration whose width disagrees with the C is one too rather than a stub that traps mid-run.
Assembly
Some of what a machine can do has no C name. An asm fn is a run of instructions given the shape of a function: the same widths cross as for a C call, because the question is the same one — what fits in a register the way the machine expects to find it.
asm fn cli() = "cli" asm fn read_cr3() -> U64 = "mov %cr3, $0" : "=r" asm fn outb(port: U16, v: U8) = "outb $1, $0" : "{dx},{al}" asm fn barrier() = "dmb sy" : "~{memory}"
The template and the constraint string go to LLVM as written, the way extern's library string goes to the linker as written: what the machine accepts is the machine's business, and a declaration is true of one architecture rather than of the language.
Leaving the constraints off asks for the ordinary shape — =r for a result and r for each parameter, in order — which is what most instructions want. Writing them out is for a register the instruction names itself: outb reads dx and al and nothing else will do. = and + open an output, ~ a clobber, and anything else is an input; a string that does not describe the signature is refused with the line it is on, because LLVM's own answer to one is to abort the compiler.
Every asm fn is emitted with side effects assumed. Two calls that look alike stay two calls, and one whose result nothing reads still runs: a device register read twice was meant twice, and cli is not there for its value. There is no way to ask for the other behaviour yet.
A function written in assembly says naked, and then the instructions are the whole of it: no prologue saving what a caller expected saved, no epilogue returning anywhere. This is what an interrupt handler, a reset vector and a context switch are.
naked asm fn isr_timer() = "push %rax\n ...\n pop %rax\n iretq" naked asm fn _kstart() = "mov $0x1000, %rsp\n call kmain\n hlt"
It takes nothing, answers nothing and has no constraints, because there is no calling convention left to carry any of them — what enters it is an interrupt or a jump somebody else wrote. Nothing in Rill may call one, and the error says so. What a program can use is the address, and that is an ordinary instruction:
asm fn isr_addr() -> Ptr = "leaq isr_timer(%rip), $0" : "=r"
Inside a naked function $ is the literal the assembler wants rather than an operand reference, so an immediate is written $0x1000, the way the manual has it.
On wasm there is no such thing: a module is verified before it runs and has no place for a machine's own encoding, so an asm fn is refused when building for one.