Standard library · Text
text/str
Imported as import "text/str" as str, its names are then str.…. Every signature below is the one the checker infers.
Strings, past what the prelude gives.
The prelude has the pieces a program cannot do without — contains, split, trim, join, the character layer. This is the rest: the searches that answer a position rather than a yes, the transforms, the padding, the parsers that refuse instead of returning zero, a hash, an edit distance, every encoding a text arrives in, and at the bottom a handful of functions that hand a long text to several strands at once.
Everything here is bytes, the way #s and s[i] are bytes: a position is a byte offset, a "character" argument is a byte value such as 32 or s[0], and upper folds ASCII letters and leaves every other byte alone. For text that needs its characters counted there is the prelude's char_* layer, and rev_chars here shows how the two combine.
Every search rides on the runtime's str_find_from and str_find_in, which cross the text sixteen bytes at a time and look at a position only where the pattern's first and last bytes both sit, so count, replace, last_index_of and the strand functions all run at memory speed. The loops that must look at every byte — upper, hash, edit_distance — are written as plain fors so the compiler lays them out as the C loop would be; the scans that stop on the data rather than on a count are whiles, whose state is written where the loop is.
str.count("the cat and the hat", "the") # => 2 str.replace("a\tb", "\t", " ") # => a b str.pad_left(int_to_str(42), 6, 32) # => 42 str.parse_int("4x2") # => None str.utf8_lossy(("ok" + chr(255) + "ok")) # => ok�ok
Functions
fn is_digit(c: Int) -> Bool
str.is_digit(48) # => true str.is_digit("a"[0]) # => false
fn is_upper(c: Int) -> Bool
str.is_upper("A"[0]) # => true
fn is_lower(c: Int) -> Bool
str.is_lower("A"[0]) # => false
fn is_alpha(c: Int) -> Bool
str.is_alpha("z"[0]) # => true str.is_alpha("9"[0]) # => false
fn is_alnum(c: Int) -> Bool
str.is_alnum("9"[0]) # => true str.is_alnum("-"[0]) # => false
fn upper_byte(c: Int) -> Int
str.upper_byte("a"[0]) # => 65
fn lower_byte(c: Int) -> Int
str.lower_byte("A"[0]) # => 97
fn has_byte(set: Str, c: Int) -> Bool
Is byte c one of the bytes of set?
str.has_byte("aeiou", "e"[0]) # => true str.has_byte("aeiou", "x"[0]) # => false
fn all_bytes(s: Str, p: (Int) -> Bool) -> Bool
Does p hold for every byte of s?
str.all_bytes("2024", str.is_digit) # => true str.all_bytes("20x4", str.is_digit) # => false
fn index_of(s: Str, p: Str) -> Int
Where p first begins, or -1. index_from starts partway along and index_in also stops early: a match must end by to.
str.index_of("hello", "ll") # => 2 str.index_of("hello", "z") # => -1
fn index_from(s: Str, p: Str, from: Int) -> Int
str.index_from("a-b-c", "-", 2) # => 3
fn index_in(s: Str, p: Str, from: Int, to: Int) -> Int
str.index_in("a-b-c", "-", 0, 2) # => 1 str.index_in("a-b-c", "-", 2, 3) # => -1
fn last_index_of(s: Str, p: Str) -> Int
Where p last begins, or -1. Found by running forward from match to match: each hop is one vectorized search, which beats walking backward a byte at a time even when the answer is near the end.
str.last_index_of("a-b-c", "-") # => 3
fn byte_index(s: Str, c: Int) -> Int
str.byte_index("hello", "l"[0]) # => 2
fn last_byte_index(s: Str, c: Int) -> Int
str.last_byte_index("hello", "l"[0]) # => 3
fn matches_at(s: Str, i: Int, p: Str) -> Bool
Does p sit at byte i of s?
str.matches_at("hello", 2, "ll") # => true str.matches_at("hello", 1, "ll") # => false
fn ends_with(s: Str, p: Str) -> Bool
str.ends_with("photo.jpg", ".jpg") # => true
fn count(s: Str, p: Str) -> Int
Occurrences of p: count skips past each match it finds, so "aa" occurs once in "aaa"; count_all steps one byte and finds it twice. Both are 0 for the empty pattern.
str.count("aaa", "aa") # => 1 str.count("the cat, the hat", "the") # => 2
fn count_all(s: Str, p: Str) -> Int
str.count_all("aaa", "aa") # => 2
fn count_by(s: Str, p: Str, step: Int) -> Int
str.count_by("aaaa", "aa", 2) # => 2
fn count_byte(s: Str, c: Int) -> Int
str.count_byte("banana", "a"[0]) # => 3
fn indices(s: Str, p: Str) -> List(Int)
Every place a non-overlapping match begins, first to last.
str.indices("a,b,c", ",") # => Cons(1, Cons(3, Nil))
fn upper(s: Str) -> Str
The byte-for-byte transforms fill a buffer of the final size and hand it over as the string, not a builder: a builder is a length store per byte, and a loop that stores a length cannot be vectorized. Written like this the loop is the one C writes, sixteen bytes at a time, and str_from_buf on a buffer nothing names again is the block itself, uncopied.
str.upper("hello, wörld") # => HELLO, WöRLD
fn lower(s: Str) -> Str
str.lower("HeLLo") # => hello
fn swap_case(s: Str) -> Str
str.swap_case("Hello") # => hELLO
fn capitalize(s: Str) -> Str
str.capitalize("hello") # => Hello
fn rev(s: Str) -> Str
Bytes back to front. On multi-byte text use rev_chars, which turns the characters around and keeps each one whole.
str.rev("abc") # => cba
fn rev_chars(s: Str) -> Str
str.rev_chars("günaydın") # => nıdyanüg
fn replace(s: Str, old: Str, new: Str) -> Str
Every occurrence of old becomes new. A string that has no old in it is handed back as it is, not copied; an empty old matches nothing. Each piece between matches goes straight from s into the builder, which is sized for the whole text up front.
str.replace("a-b-c", "-", "+") # => a+b+c str.replace("abc", "", "x") # => abc
fn replace_first(s: Str, old: Str, new: Str) -> Str
str.replace_first("a-b-c", "-", "+") # => a+b-c
fn pad_left(s: Str, n: Int, c: Int) -> Str
Bring s up to n bytes with byte c; a string already that long is returned as it is.
str.pad_left("42", 5, "0"[0]) # => 00042 str.pad_left("hello", 3, 32) # => hello
fn pad_right(s: Str, n: Int, c: Int) -> Str
str.pad_right("ab", 4, "."[0]) # => ab..
fn center(s: Str, n: Int, c: Int) -> Str
str.center("ab", 6, "*"[0]) # => **ab**
fn trim_left(s: Str) -> Str
str.trim_left(" a ") # => a
fn trim_right(s: Str) -> Str
str.trim_right(" a ") # => a
fn trim_set(s: Str, set: Str) -> Str
Trim any of the bytes in set from both ends: trim_set(path, "/").
str.trim_set("/usr/local/", "/") # => usr/local
fn strip_prefix(s: Str, p: Str) -> Option(Str)
The rest of s once p has been taken off its front (or its back), or None when p was not there — so a caller can tell "had no prefix" from "had the prefix and nothing else".
str.strip_prefix("v0.13", "v") # => Some(0.13) str.strip_prefix("0.13", "v") # => None
fn strip_suffix(s: Str, p: Str) -> Option(Str)
str.strip_suffix("main.rill", ".rill") # => Some(main)
fn words(s: Str) -> List(Str)
The runs of non-space bytes, in order; no empty strings, however many spaces there are.
str.words(" two words ") # => Cons(two, Cons(words, Nil))
fn skip_ws(s: Str, i: Int) -> Int
str.skip_ws(" x", 0) # => 3
fn split_once(s: Str, sep: Str) -> Option((Str, Str))
The text before and after the first sep, or None when there is none: split_once(header, ": ").
str.split_once("Host: a.b", ": ") # => Some((Host, a.b)) str.split_once("none", ":") # => None
fn split_at(s: Str, i: Int) -> (Str, Str)
str.split_at("abcdef", 2) # => (ab, cdef)
fn concat(l: List(Str)) -> Str
str.concat(Cons("a", Cons("b", Cons("c", Nil)))) # => abc
fn line_count(s: Str) -> Int
What lines would return the length of, without building it.
str.line_count("one\ntwo\n") # => 2
fn cmp(a: Str, b: Str) -> Int
str.cmp("apple", "banana") # => -1 str.cmp("b", "b") # => 0
fn eq_ignore_case(a: Str, b: Str) -> Bool
str.eq_ignore_case("Content-Type", "content-type") # => true
fn common_prefix(a: Str, b: Str) -> Int
Length of the longest prefix the two share.
str.common_prefix("interstellar", "internet") # => 5
fn hash(s: Str) -> Int
FNV-1a over the bytes, 64 bits wide. The multiply is meant to wrap.
str.hash("a") == str.hash("a") # => true str.hash("a") == str.hash("b") # => false
fn edit_distance(a: Str, b: Str) -> Int
Levenshtein distance: the fewest single-byte edits (insert, delete, replace) that turn a into b. Two rows of the usual table, swapped rather than copied, so it is #b + 1 words of memory however long a is.
str.edit_distance("kitten", "sitting") # => 3
fn parse_int(s: Str) -> Option(Int)
A whole decimal integer, with an optional sign, or None. This is the strict reading str_to_int does not offer: "12ab" and "" are None here rather than 12 and 0.
str.parse_int("-42") # => Some(-42) str.parse_int("12ab") # => None str.parse_int("") # => None
fn digits(s: Str, i: Int) -> Option(Int)
The digits from i to the end of s, as a number; None if anything else is there, or nothing is.
str.digits("id-1234", 3) # => Some(1234) str.digits("id-12x4", 3) # => None
fn is_int(s: Str) -> Bool
str.is_int("-7") # => true str.is_int("7.0") # => false
fn is_blank(s: Str) -> Bool
str.is_blank(" \t") # => true str.is_blank(" a ") # => false
fn is_alpha_str(s: Str) -> Bool
str.is_alpha_str("abc") # => true str.is_alpha_str("ab1") # => false
fn is_digit_str(s: Str) -> Bool
str.is_digit_str("2024") # => true
fn bytes(s: Str) -> Buf(U8)
The bytes as a buffer, and back.
#str.bytes("héllo") # => 6
fn from_bytes(v: Buf(U8)) -> Str
str.from_bytes(str.bytes("round trip")) # => round trip
fn utf8_seq_len(s: Str, i: Int) -> Int
Is the byte at i the start of a well-formed sequence, and how long is it? 1 to 4, or 0 for a stray continuation byte, an overlong form, a surrogate, a code point past U+10FFFF, or a sequence the string ends inside. Reads past the end are -1 and fail every test, so nothing here looks at #s.
str.utf8_seq_len("é", 0) # => 2 str.utf8_seq_len("a", 0) # => 1 str.utf8_seq_len("é", 1) # => 0
fn is_cont(b: Int) -> Bool
str.is_cont("é"[1]) # => true str.is_cont("a"[0]) # => false
fn utf8_bad_at(s: Str) -> Int
The offset of the first byte that is not UTF-8, or -1 when all of it is.
str.utf8_bad_at(("ok" + chr(255) + "ok")) # => 2 str.utf8_bad_at("all fine") # => -1
fn utf8_valid(s: Str) -> Bool
str.utf8_valid("günaydın") # => true str.utf8_valid(chr(195)) # => false
fn is_scalar(c: Int) -> Bool
A code point that UTF-8 may encode: not a surrogate, not past the last plane.
str.is_scalar(233) # => true str.is_scalar(55296) # => false
fn decode_utf8(s: Str) -> Option(List(Int))
str.decode_utf8("aé") # => Some(Cons(97, Cons(233, Nil))) str.decode_utf8(chr(255)) # => None
fn encode_utf8(cps: List(Int)) -> Option(Str)
str.encode_utf8(Cons(97, Cons(233, Nil))) # => Some(aé)
fn utf8_lossy(s: Str) -> Str
The text with every byte that is not part of a well-formed sequence replaced by U+FFFD, one replacement per bad byte. Valid text comes back as it is, uncopied.
str.utf8_lossy(("ok" + chr(255) + "ok")) # => ok�ok
fn has_bom(s: Str) -> Bool
str.has_bom((chr(239) + chr(187) + chr(191) + "x")) # => true str.has_bom("x") # => false
fn strip_bom(s: Str) -> Str
str.strip_bom((chr(239) + chr(187) + chr(191) + "x")) # => x
fn is_ascii(s: Str) -> Bool
str.is_ascii("plain") # => true str.is_ascii("café") # => false
fn each_char(s: Str, i: Int, f: (Int, Int) -> 'a) -> Int
Walk the characters of s from byte i, giving f each code point and where it starts.
str.each_char("aé", 0, \c, at -> c) # => 3
fn map_chars(s: Str, f: (Int) -> Int) -> Str
The text with f applied to every code point, built without a list.
str.map_chars("abc", \c -> c + 1) # => bcd
fn put_char(b: StrBuf, c: Int) -> Unit
char_onto from the prelude, but a statement: it answers nothing, so it can sit in a branch whose other arms are b += x.
b = strbuf() str.put_char(b, 233) strbuf_str(b) # => é
fn encode_utf16(s: Str, big: Bool) -> Str
str.hex_encode(str.encode_utf16("A", true)) # => 0041
fn encode_utf16le(s: Str) -> Str
str.hex_encode(str.encode_utf16le("A")) # => 4100
fn encode_utf16be(s: Str) -> Str
str.hex_encode(str.encode_utf16be("é")) # => 00e9
fn decode_utf16(v: Str, big: Bool) -> Option(Str)
str.decode_utf16((chr(0) + "A"), true) # => Some(A)
fn decode_utf16le(v: Str) -> Option(Str)
str.decode_utf16le(("A" + chr(0))) # => Some(A)
fn decode_utf16be(v: Str) -> Option(Str)
str.decode_utf16be((chr(0) + chr(233))) # => Some(é) str.decode_utf16be(chr(0)) # => None
fn decode_utf16_bom(v: Str) -> Option(Str)
str.decode_utf16_bom((chr(255) + chr(254) + "A" + chr(0))) # => Some(A)
fn encode_utf32(s: Str, big: Bool) -> Str
str.hex_encode(str.encode_utf32("A", true)) # => 00000041
fn decode_utf32(v: Str, big: Bool) -> Option(Str)
str.decode_utf32((chr(0) + chr(0) + chr(0) + "A"), true) # => Some(A)
fn decode_bytes(v: Str, f: (Int) -> Int) -> Str
str.decode_bytes(chr(233), str.cp1252_char) # => é
fn encode_bytes(s: Str, g: (Int) -> Int) -> Option(Str)
str.hex_encode(str.encode_bytes("é", \c -> if c < 256 then c else -1) |> unwrap_or("")) # => e9
fn decode_latin1(v: Str) -> Str
str.decode_latin1(("caf" + chr(233))) # => café
fn encode_latin1(s: Str) -> Option(Str)
str.hex_encode(str.encode_latin1("é") |> unwrap_or("")) # => e9 str.encode_latin1("ş") # => None
fn cp1252_char(b: Int) -> Int
Windows-1252: Latin-1 with the C1 controls replaced by punctuation and a few letters. The five bytes 1252 leaves undefined decode to themselves.
str.cp1252_char(128) # => 8364 str.cp1252_char(233) # => 233
fn cp1254_char(b: Int) -> Int
Windows-1254, the Turkish page: 1252 with Ğ İ Ş ğ ı ş where 1252 has Ð Ý Þ ð ý þ, and without Ž ž.
str.cp1254_char(208) # => 286
fn decode_cp1252(v: Str) -> Str
str.decode_cp1252(chr(128)) # => €
fn encode_cp1252(s: Str) -> Option(Str)
str.hex_encode(str.encode_cp1252("€") |> unwrap_or("")) # => 80
fn decode_cp1254(v: Str) -> Str
str.decode_cp1254(chr(208)) # => Ğ
fn encode_cp1254(s: Str) -> Option(Str)
str.hex_encode(str.encode_cp1254("Ğ") |> unwrap_or("")) # => d0
fn char_upper(c: Int) -> Int
str.char_upper(233) # => 201
fn char_lower(c: Int) -> Int
str.char_lower(201) # => 233
fn char_upper_tr(c: Int) -> Int
str.char_upper_tr("i"[0]) # => 304
fn char_lower_tr(c: Int) -> Int
str.char_lower_tr("I"[0]) # => 305
fn upper_utf8(s: Str) -> Str
str.upper_utf8("café") # => CAFÉ
fn lower_utf8(s: Str) -> Str
str.lower_utf8("CAFÉ") # => café
fn upper_tr(s: Str) -> Str
str.upper_tr("istanbul") # => İSTANBUL
fn lower_tr(s: Str) -> Str
str.lower_tr("ISPARTA") # => ısparta
fn ascii_fold(s: Str) -> Str
The nearest ASCII letter for an accented Latin one — ç→c, ğ→g, ı→i, İ→I, ö→o, ş→s, ü→u, é→e and the rest of Latin-1 and Extended-A — with ß, æ and œ opening out to two letters. Everything else is kept. For slugs, file names and the search box that must find "Çağrı" when "cagri" is typed.
str.ascii_fold("Çağrı Straße") # => Cagri Strasse
fn base64_encode(s: Str) -> Str
str.base64_encode("hello") # => aGVsbG8=
fn base64url_encode(s: Str) -> Str
The URL-safe alphabet, without padding, as JWTs and the like carry it.
str.base64url_encode((chr(251) + chr(255))) # => -_8
fn base64_with(s: Str, abc: Str, pad: Bool) -> Str
str.base64_with("hi", "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/", false) # => aGk
fn base64_decode(s: Str) -> Option(Str)
Decodes either alphabet, with or without padding. Anything else in the input — a space, a newline, a third alphabet — is None.
str.base64_decode("aGVsbG8=") # => Some(hello) str.base64_decode("aGVs bG8=") # => None
fn base64url_decode(s: Str) -> Option(Str)
str.hex_encode(str.base64url_decode("-_8") |> unwrap_or("")) # => fbff
fn hex_encode(s: Str) -> Str
str.hex_encode("AZ") # => 415a
fn hex_decode(s: Str) -> Option(Str)
str.hex_decode("415a") # => Some(AZ) str.hex_decode("41g") # => None
fn int_to_hex(n: Int) -> Str
str.int_to_hex(255) # => ff
fn hex_to_int(s: Str) -> Option(Int)
"ff", "0xFF" and "-0x1a" all read; anything else is None.
str.hex_to_int("0xFF") # => Some(255) str.hex_to_int("-0x1a") # => Some(-26) str.hex_to_int("zz") # => None
fn hex_digits(s: Str, i: Int) -> Option(Int)
The hex digits from i to the end of s, as a number; None if anything else is there.
str.hex_digits("#ff", 1) # => Some(255)
fn url_encode(s: Str) -> Str
Percent-encoding as RFC 3986 has it: letters, digits and -_.~ stand, every other byte becomes %XX. form_encode is the variant HTML forms use, with a space as +.
str.url_encode("a b/c") # => a%20b%2Fc
fn form_encode(s: Str) -> Str
str.form_encode("a b&c") # => a+b%26c
fn url_decode(s: Str) -> Option(Str)
A % not followed by two hex digits is None.
str.url_decode("a%20b") # => Some(a b) str.url_decode("a%2") # => None
fn form_decode(s: Str) -> Option(Str)
str.form_decode("a+b%26c") # => Some(a b&c)
fn html_escape(s: Str) -> Str
The five characters HTML cannot show as they are. Safe for text and for attribute values in either kind of quote.
str.html_escape("<a href=\"x\">") # => <a href="x">
fn html_unescape(s: Str) -> Str
The five named entities and the numeric ones, decimal and hex. An & that starts no entity is kept as it is, which is what browsers do.
str.html_unescape("<b> & é A") # => <b> & é A
fn escape(s: Str) -> Str
C-style escapes: \n \r \t \\ \" and \xHH for any other control byte, so that a string can be shown on one line and read back exactly.
str.escape("a\tb\n") # => a\tb\n
fn unescape(s: Str) -> Option(Str)
Reads escape's output back, plus \', \0 and \u{...}. A backslash that starts no escape is None.
str.unescape("a\\tb") # => Some(a b) str.unescape("\\u{e9}") # => Some(é) str.unescape("\\q") # => None
fn par_count(s: Str, p: Str) -> Int
count_all (every occurrence, overlapping ones included) over all the workers. The overlapping count is the one that splits: whether a non-overlapping scan takes a match depends on the matches before it, which another strand may hold. For a pattern that cannot overlap itself the two counts are the same number.
str.par_count("aaa", "aa") # => 2
fn par_count_on(s: Str, p: Str, k: Int) -> Int
str.par_count_on("the cat, the hat", "the", 4) # => 2
fn par_index_of(s: Str, p: Str) -> Int
index_of over all the workers: each range reports the first match that begins in it, and the smallest wins.
str.par_index_of("hello", "ll") # => 2
fn par_index_on(s: Str, p: Str, k: Int) -> Int
str.par_index_on("hello", "lo", 3) # => 3
fn par_lines(s: Str, f: (Str) -> 'a) -> List('a)
map(lines(s), f) over all the workers: the ranges are moved forward to the byte after a newline so that no line is cut, each strand maps its own lines, and the pieces are joined in order. The lines are exactly those of the prelude's lines, trailing empty one included.
str.par_lines("a\nbb\n", \l -> #l) # => Cons(1, Cons(2, Cons(0, Nil)))
fn par_lines_on(s: Str, f: (Str) -> 'a, k: Int) -> List('a)
str.par_lines_on("a\nbb\nccc", \l -> #l, 2) # => Cons(1, Cons(2, Cons(3, Nil)))