Chapter 10
Files
read_file and write_file are for a file a program wants all of at once. This section is for the other kind: a file too big to hold, or one being appended to for as long as the program runs β a log, a journal, a store whose readers are at different places in it.
A file is a descriptor, the same Int a socket is, and for the same reason: the two families are then written alike, and a program holding both is not holding two ideas of what an open thing is.
fn append(path, record) = fd = file_open(path, "a") at = file_size(fd) file_write(fd, record) file_sync(fd) file_close(fd) at
file_open(path, mode) -> Int |
a descriptor, or -1. "r" "w" "a", and each with + for reading as well: "w" truncates or creates, "a" appends and creates |
file_read(fd, n) -> Str |
up to n bytes from where the descriptor is; "" at the end |
file_pread(fd, off, n) -> Str |
up to n bytes from off, leaving the descriptor where it was |
file_pread_into(fd, off, buf, at, n) -> Int |
the same, into n bytes of a Buf(U8) from at, with no string made: how many arrived, or -1 |
file_write(fd, s) -> Int |
writes all of it; how many bytes went, or -1 |
file_pwrite(fd, off, buf, at, n) -> Int |
n bytes of a Buf(U8) from at to offset off, leaving the descriptor where it was; all of it, or -1 |
file_lock(fd, exclusive) -> Bool file_unlock(fd) -> Bool |
the whole-file advisory lock (flock), taken at once or not at all β a lock another process holds is false, not a wait |
file_seek(fd, off) -> Int |
move to an absolute offset |
file_pos(fd) -> Int file_size(fd) -> Int |
where it is, and how big the file is |
file_truncate(fd, n) -> Bool |
cut it to n bytes, or grow it with zeros |
file_sync(fd) -> Bool |
wait until what has been written is on the disk |
file_close(fd) |
|
file_remove(path) -> Bool file_rename(from, to) -> Bool file_exists(path) -> Bool |
|
dir_make(path) -> Bool dir_list(path) -> Str dir_sync(path) -> Bool |
a directory: make one, its names one to a line, and make those names durable |
The prelude has open_file, which is file_open as a Result the way listen_on is tcp_listen as one; dir_names, which is dir_list as a list; and replace_file, which is the four steps below written once.
file_pread_into and file_pwrite are for a page cache. A store that keeps its pages in one buffer and reads the same-sized block into a slot of it wants no string made on the way, and writes a page back where it came from without moving the descriptor. lib/db/kv.rill is written on the two of them, and on file_lock, which is how it keeps a second process out.
file_pread is what lets one file answer several readers. A consumer reading a log from where it left off does not move the descriptor, so it does not disturb another consumer at a different place in the file, or the producer appending to the end of it. A program that reads with file_read is a program whose readers have to take turns.
A write that has returned is not on the disk. It is in the operating system's cache, and a machine that loses power between the two loses it. file_sync is the wait for the disk to say it has the bytes, and it is the only file call slow enough β a millisecond on a good drive, ten on a bad one β to be worth thinking about. On a Mac it is F_FULLFSYNC rather than fsync, because fsync there hands the bytes to the drive and returns without waiting for the drive to keep them, which is fast and is not durability.
A sync does not hold the worker. This is the one place in the section where the concurrency model shows through. Every other file call is a syscall that returns; a sync hands the work to a helper thread and parks the strand on a pipe, which is the same manoeuvre a strand waiting on a socket makes and goes through the same poller. The worker is free from the moment the request is handed over. A strand computing beside a strand syncing takes about as long as it takes on its own, which is what a_sync_does_not_hold_the_worker measures. Before the scheduler exists β a program that neither spawns nor waits runs main on the process stack β the sync is simply done where it stands, because there is no worker to spare and nothing to be an improvement on.
Syncing a file does not make its name durable. Creating a file and renaming one both change the directory, and that change is cached like any other: a program that creates a file, syncs it, and loses power can come back to a directory that has never heard of it. dir_sync is the call that closes that gap. It is separate so that a program pays for it after a create or a rename rather than after every write.
Those two together are how a file is replaced without a moment in which it is neither the old one nor the new one β write a temporary, sync it, rename it over, sync the directory. A rename on a Unix filesystem is atomic, so a reader sees one file or the other and never half of either. replace_file in the prelude is those four steps; examples/files.rill is the rest of this section.
Nothing here buffers. A write is a write, so a program appending a hundred small records makes a hundred syscalls; build the bytes in a StrBuf and write once, which is what the language would have you do anyway.