Rill v0.13 Reference

Standard library · Text

text/str

Imported as import "text/str" as str, its names are then str.…. Every signature below is the one the checker infers.

Strings, past what the prelude gives.

The prelude has the pieces a program cannot do without — contains, split, trim, join, the character layer. This is the rest: the searches that answer a position rather than a yes, the transforms, the padding, the parsers that refuse instead of returning zero, a hash, an edit distance, every encoding a text arrives in, and at the bottom a handful of functions that hand a long text to several strands at once.

Everything here is bytes, the way #s and s[i] are bytes: a position is a byte offset, a "character" argument is a byte value such as 32 or s[0], and upper folds ASCII letters and leaves every other byte alone. For text that needs its characters counted there is the prelude's char_* layer, and rev_chars here shows how the two combine.

Every search rides on the runtime's str_find_from and str_find_in, which cross the text sixteen bytes at a time and look at a position only where the pattern's first and last bytes both sit, so count, replace, last_index_of and the strand functions all run at memory speed. The loops that must look at every byte — upper, hash, edit_distance — are written as plain fors so the compiler lays them out as the C loop would be; the scans that stop on the data rather than on a count are whiles, whose state is written where the loop is.

str.count("the cat and the hat", "the")        # => 2
str.replace("a\tb", "\t", "    ")              # => a    b
str.pad_left(int_to_str(42), 6, 32)              # =>     42
str.parse_int("4x2")                           # => None
str.utf8_lossy(("ok" + chr(255) + "ok"))                     # => ok�ok

Functions

fn is_digit(c: Int) -> Bool

str.is_digit(48)         # => true
str.is_digit("a"[0])     # => false

fn is_upper(c: Int) -> Bool

str.is_upper("A"[0])     # => true

fn is_lower(c: Int) -> Bool

str.is_lower("A"[0])     # => false

fn is_alpha(c: Int) -> Bool

str.is_alpha("z"[0])     # => true
str.is_alpha("9"[0])     # => false

fn is_alnum(c: Int) -> Bool

str.is_alnum("9"[0])     # => true
str.is_alnum("-"[0])     # => false

fn upper_byte(c: Int) -> Int

str.upper_byte("a"[0])   # => 65

fn lower_byte(c: Int) -> Int

str.lower_byte("A"[0])   # => 97

fn has_byte(set: Str, c: Int) -> Bool

Is byte c one of the bytes of set?

str.has_byte("aeiou", "e"[0])   # => true
str.has_byte("aeiou", "x"[0])   # => false

fn all_bytes(s: Str, p: (Int) -> Bool) -> Bool

Does p hold for every byte of s?

str.all_bytes("2024", str.is_digit)   # => true
str.all_bytes("20x4", str.is_digit)   # => false

fn index_of(s: Str, p: Str) -> Int

Where p first begins, or -1. index_from starts partway along and index_in also stops early: a match must end by to.

str.index_of("hello", "ll")    # => 2
str.index_of("hello", "z")     # => -1

fn index_from(s: Str, p: Str, from: Int) -> Int

str.index_from("a-b-c", "-", 2)   # => 3

fn index_in(s: Str, p: Str, from: Int, to: Int) -> Int

str.index_in("a-b-c", "-", 0, 2)   # => 1
str.index_in("a-b-c", "-", 2, 3)   # => -1

fn last_index_of(s: Str, p: Str) -> Int

Where p last begins, or -1. Found by running forward from match to match: each hop is one vectorized search, which beats walking backward a byte at a time even when the answer is near the end.

str.last_index_of("a-b-c", "-")   # => 3

fn byte_index(s: Str, c: Int) -> Int

str.byte_index("hello", "l"[0])   # => 2

fn last_byte_index(s: Str, c: Int) -> Int

str.last_byte_index("hello", "l"[0])   # => 3

fn matches_at(s: Str, i: Int, p: Str) -> Bool

Does p sit at byte i of s?

str.matches_at("hello", 2, "ll")   # => true
str.matches_at("hello", 1, "ll")   # => false

fn ends_with(s: Str, p: Str) -> Bool

str.ends_with("photo.jpg", ".jpg")   # => true

fn count(s: Str, p: Str) -> Int

Occurrences of p: count skips past each match it finds, so "aa" occurs once in "aaa"; count_all steps one byte and finds it twice. Both are 0 for the empty pattern.

str.count("aaa", "aa")          # => 1
str.count("the cat, the hat", "the")   # => 2

fn count_all(s: Str, p: Str) -> Int

str.count_all("aaa", "aa")      # => 2

fn count_by(s: Str, p: Str, step: Int) -> Int

str.count_by("aaaa", "aa", 2)    # => 2

fn count_byte(s: Str, c: Int) -> Int

str.count_byte("banana", "a"[0])   # => 3

fn indices(s: Str, p: Str) -> List(Int)

Every place a non-overlapping match begins, first to last.

str.indices("a,b,c", ",")   # => Cons(1, Cons(3, Nil))

fn upper(s: Str) -> Str

The byte-for-byte transforms fill a buffer of the final size and hand it over as the string, not a builder: a builder is a length store per byte, and a loop that stores a length cannot be vectorized. Written like this the loop is the one C writes, sixteen bytes at a time, and str_from_buf on a buffer nothing names again is the block itself, uncopied.

str.upper("hello, wörld")   # => HELLO, WöRLD

fn lower(s: Str) -> Str

str.lower("HeLLo")   # => hello

fn swap_case(s: Str) -> Str

str.swap_case("Hello")   # => hELLO

fn capitalize(s: Str) -> Str

str.capitalize("hello")   # => Hello

fn rev(s: Str) -> Str

Bytes back to front. On multi-byte text use rev_chars, which turns the characters around and keeps each one whole.

str.rev("abc")   # => cba

fn rev_chars(s: Str) -> Str

str.rev_chars("günaydın")   # => nıdyanüg

fn replace(s: Str, old: Str, new: Str) -> Str

Every occurrence of old becomes new. A string that has no old in it is handed back as it is, not copied; an empty old matches nothing. Each piece between matches goes straight from s into the builder, which is sized for the whole text up front.

str.replace("a-b-c", "-", "+")   # => a+b+c
str.replace("abc", "", "x")     # => abc

fn replace_first(s: Str, old: Str, new: Str) -> Str

str.replace_first("a-b-c", "-", "+")   # => a+b-c

fn pad_left(s: Str, n: Int, c: Int) -> Str

Bring s up to n bytes with byte c; a string already that long is returned as it is.

str.pad_left("42", 5, "0"[0])   # => 00042
str.pad_left("hello", 3, 32)   # => hello

fn pad_right(s: Str, n: Int, c: Int) -> Str

str.pad_right("ab", 4, "."[0])   # => ab..

fn center(s: Str, n: Int, c: Int) -> Str

str.center("ab", 6, "*"[0])   # => **ab**

fn trim_left(s: Str) -> Str

str.trim_left("  a  ")   # => a

fn trim_right(s: Str) -> Str

str.trim_right("  a  ")   # =>   a

fn trim_set(s: Str, set: Str) -> Str

Trim any of the bytes in set from both ends: trim_set(path, "/").

str.trim_set("/usr/local/", "/")   # => usr/local

fn strip_prefix(s: Str, p: Str) -> Option(Str)

The rest of s once p has been taken off its front (or its back), or None when p was not there — so a caller can tell "had no prefix" from "had the prefix and nothing else".

str.strip_prefix("v0.13", "v")   # => Some(0.13)
str.strip_prefix("0.13", "v")    # => None

fn strip_suffix(s: Str, p: Str) -> Option(Str)

str.strip_suffix("main.rill", ".rill")   # => Some(main)

fn words(s: Str) -> List(Str)

The runs of non-space bytes, in order; no empty strings, however many spaces there are.

str.words("  two   words ")   # => Cons(two, Cons(words, Nil))

fn skip_ws(s: Str, i: Int) -> Int

str.skip_ws("   x", 0)   # => 3

fn split_once(s: Str, sep: Str) -> Option((Str, Str))

The text before and after the first sep, or None when there is none: split_once(header, ": ").

str.split_once("Host: a.b", ": ")   # => Some((Host, a.b))
str.split_once("none", ":")        # => None

fn split_at(s: Str, i: Int) -> (Str, Str)

str.split_at("abcdef", 2)   # => (ab, cdef)

fn concat(l: List(Str)) -> Str

str.concat(Cons("a", Cons("b", Cons("c", Nil))))   # => abc

fn line_count(s: Str) -> Int

What lines would return the length of, without building it.

str.line_count("one\ntwo\n")   # => 2

fn cmp(a: Str, b: Str) -> Int

str.cmp("apple", "banana")   # => -1
str.cmp("b", "b")            # => 0

fn eq_ignore_case(a: Str, b: Str) -> Bool

str.eq_ignore_case("Content-Type", "content-type")   # => true

fn common_prefix(a: Str, b: Str) -> Int

Length of the longest prefix the two share.

str.common_prefix("interstellar", "internet")   # => 5

fn hash(s: Str) -> Int

FNV-1a over the bytes, 64 bits wide. The multiply is meant to wrap.

str.hash("a") == str.hash("a")   # => true
str.hash("a") == str.hash("b")   # => false

fn edit_distance(a: Str, b: Str) -> Int

Levenshtein distance: the fewest single-byte edits (insert, delete, replace) that turn a into b. Two rows of the usual table, swapped rather than copied, so it is #b + 1 words of memory however long a is.

str.edit_distance("kitten", "sitting")   # => 3

fn parse_int(s: Str) -> Option(Int)

A whole decimal integer, with an optional sign, or None. This is the strict reading str_to_int does not offer: "12ab" and "" are None here rather than 12 and 0.

str.parse_int("-42")    # => Some(-42)
str.parse_int("12ab")   # => None
str.parse_int("")       # => None

fn digits(s: Str, i: Int) -> Option(Int)

The digits from i to the end of s, as a number; None if anything else is there, or nothing is.

str.digits("id-1234", 3)   # => Some(1234)
str.digits("id-12x4", 3)   # => None

fn is_int(s: Str) -> Bool

str.is_int("-7")   # => true
str.is_int("7.0")  # => false

fn is_blank(s: Str) -> Bool

str.is_blank("  \t")   # => true
str.is_blank(" a ")    # => false

fn is_alpha_str(s: Str) -> Bool

str.is_alpha_str("abc")   # => true
str.is_alpha_str("ab1")   # => false

fn is_digit_str(s: Str) -> Bool

str.is_digit_str("2024")   # => true

fn bytes(s: Str) -> Buf(U8)

The bytes as a buffer, and back.

#str.bytes("héllo")   # => 6

fn from_bytes(v: Buf(U8)) -> Str

str.from_bytes(str.bytes("round trip"))   # => round trip

fn utf8_seq_len(s: Str, i: Int) -> Int

Is the byte at i the start of a well-formed sequence, and how long is it? 1 to 4, or 0 for a stray continuation byte, an overlong form, a surrogate, a code point past U+10FFFF, or a sequence the string ends inside. Reads past the end are -1 and fail every test, so nothing here looks at #s.

str.utf8_seq_len("é", 0)   # => 2
str.utf8_seq_len("a", 0)   # => 1
str.utf8_seq_len("é", 1)   # => 0

fn is_cont(b: Int) -> Bool

str.is_cont("é"[1])   # => true
str.is_cont("a"[0])   # => false

fn utf8_bad_at(s: Str) -> Int

The offset of the first byte that is not UTF-8, or -1 when all of it is.

str.utf8_bad_at(("ok" + chr(255) + "ok"))   # => 2
str.utf8_bad_at("all fine")   # => -1

fn utf8_valid(s: Str) -> Bool

str.utf8_valid("günaydın")   # => true
str.utf8_valid(chr(195))       # => false

fn is_scalar(c: Int) -> Bool

A code point that UTF-8 may encode: not a surrogate, not past the last plane.

str.is_scalar(233)      # => true
str.is_scalar(55296)    # => false

fn decode_utf8(s: Str) -> Option(List(Int))

str.decode_utf8("aé")   # => Some(Cons(97, Cons(233, Nil)))
str.decode_utf8(chr(255))  # => None

fn encode_utf8(cps: List(Int)) -> Option(Str)

str.encode_utf8(Cons(97, Cons(233, Nil)))   # => Some(aé)

fn utf8_lossy(s: Str) -> Str

The text with every byte that is not part of a well-formed sequence replaced by U+FFFD, one replacement per bad byte. Valid text comes back as it is, uncopied.

str.utf8_lossy(("ok" + chr(255) + "ok"))   # => ok�ok

fn has_bom(s: Str) -> Bool

str.has_bom((chr(239) + chr(187) + chr(191) + "x"))   # => true
str.has_bom("x")   # => false

fn strip_bom(s: Str) -> Str

str.strip_bom((chr(239) + chr(187) + chr(191) + "x"))   # => x

fn is_ascii(s: Str) -> Bool

str.is_ascii("plain")   # => true
str.is_ascii("café")    # => false

fn each_char(s: Str, i: Int, f: (Int, Int) -> 'a) -> Int

Walk the characters of s from byte i, giving f each code point and where it starts.

str.each_char("aé", 0, \c, at -> c)   # => 3

fn map_chars(s: Str, f: (Int) -> Int) -> Str

The text with f applied to every code point, built without a list.

str.map_chars("abc", \c -> c + 1)   # => bcd

fn put_char(b: StrBuf, c: Int) -> Unit

char_onto from the prelude, but a statement: it answers nothing, so it can sit in a branch whose other arms are b += x.

b = strbuf()
str.put_char(b, 233)
strbuf_str(b)   # => é

fn encode_utf16(s: Str, big: Bool) -> Str

str.hex_encode(str.encode_utf16("A", true))   # => 0041

fn encode_utf16le(s: Str) -> Str

str.hex_encode(str.encode_utf16le("A"))   # => 4100

fn encode_utf16be(s: Str) -> Str

str.hex_encode(str.encode_utf16be("é"))   # => 00e9

fn decode_utf16(v: Str, big: Bool) -> Option(Str)

str.decode_utf16((chr(0) + "A"), true)   # => Some(A)

fn decode_utf16le(v: Str) -> Option(Str)

str.decode_utf16le(("A" + chr(0)))   # => Some(A)

fn decode_utf16be(v: Str) -> Option(Str)

str.decode_utf16be((chr(0) + chr(233)))   # => Some(é)
str.decode_utf16be(chr(0))       # => None

fn decode_utf16_bom(v: Str) -> Option(Str)

str.decode_utf16_bom((chr(255) + chr(254) + "A" + chr(0)))   # => Some(A)

fn encode_utf32(s: Str, big: Bool) -> Str

str.hex_encode(str.encode_utf32("A", true))   # => 00000041

fn decode_utf32(v: Str, big: Bool) -> Option(Str)

str.decode_utf32((chr(0) + chr(0) + chr(0) + "A"), true)   # => Some(A)

fn decode_bytes(v: Str, f: (Int) -> Int) -> Str

str.decode_bytes(chr(233), str.cp1252_char)   # => é

fn encode_bytes(s: Str, g: (Int) -> Int) -> Option(Str)

str.hex_encode(str.encode_bytes("é", \c -> if c < 256 then c else -1) |> unwrap_or(""))   # => e9

fn decode_latin1(v: Str) -> Str

str.decode_latin1(("caf" + chr(233)))   # => café

fn encode_latin1(s: Str) -> Option(Str)

str.hex_encode(str.encode_latin1("é") |> unwrap_or(""))   # => e9
str.encode_latin1("ş")   # => None

fn cp1252_char(b: Int) -> Int

Windows-1252: Latin-1 with the C1 controls replaced by punctuation and a few letters. The five bytes 1252 leaves undefined decode to themselves.

str.cp1252_char(128)   # => 8364
str.cp1252_char(233)   # => 233

fn cp1254_char(b: Int) -> Int

Windows-1254, the Turkish page: 1252 with Ğ İ Ş ğ ı ş where 1252 has Ð Ý Þ ð ý þ, and without Ž ž.

str.cp1254_char(208)   # => 286

fn decode_cp1252(v: Str) -> Str

str.decode_cp1252(chr(128))   # => €

fn encode_cp1252(s: Str) -> Option(Str)

str.hex_encode(str.encode_cp1252("€") |> unwrap_or(""))   # => 80

fn decode_cp1254(v: Str) -> Str

str.decode_cp1254(chr(208))   # => Ğ

fn encode_cp1254(s: Str) -> Option(Str)

str.hex_encode(str.encode_cp1254("Ğ") |> unwrap_or(""))   # => d0

fn char_upper(c: Int) -> Int

str.char_upper(233)   # => 201

fn char_lower(c: Int) -> Int

str.char_lower(201)   # => 233

fn char_upper_tr(c: Int) -> Int

str.char_upper_tr("i"[0])   # => 304

fn char_lower_tr(c: Int) -> Int

str.char_lower_tr("I"[0])   # => 305

fn upper_utf8(s: Str) -> Str

str.upper_utf8("café")   # => CAFÉ

fn lower_utf8(s: Str) -> Str

str.lower_utf8("CAFÉ")   # => café

fn upper_tr(s: Str) -> Str

str.upper_tr("istanbul")   # => İSTANBUL

fn lower_tr(s: Str) -> Str

str.lower_tr("ISPARTA")   # => ısparta

fn ascii_fold(s: Str) -> Str

The nearest ASCII letter for an accented Latin one — ç→c, ğ→g, ı→i, İ→I, ö→o, ş→s, ü→u, é→e and the rest of Latin-1 and Extended-A — with ß, æ and œ opening out to two letters. Everything else is kept. For slugs, file names and the search box that must find "Çağrı" when "cagri" is typed.

str.ascii_fold("Çağrı Straße")   # => Cagri Strasse

fn base64_encode(s: Str) -> Str

str.base64_encode("hello")   # => aGVsbG8=

fn base64url_encode(s: Str) -> Str

The URL-safe alphabet, without padding, as JWTs and the like carry it.

str.base64url_encode((chr(251) + chr(255)))   # => -_8

fn base64_with(s: Str, abc: Str, pad: Bool) -> Str

str.base64_with("hi", "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/", false)   # => aGk

fn base64_decode(s: Str) -> Option(Str)

Decodes either alphabet, with or without padding. Anything else in the input — a space, a newline, a third alphabet — is None.

str.base64_decode("aGVsbG8=")   # => Some(hello)
str.base64_decode("aGVs bG8=")  # => None

fn base64url_decode(s: Str) -> Option(Str)

str.hex_encode(str.base64url_decode("-_8") |> unwrap_or(""))   # => fbff

fn hex_encode(s: Str) -> Str

str.hex_encode("AZ")   # => 415a

fn hex_decode(s: Str) -> Option(Str)

str.hex_decode("415a")   # => Some(AZ)
str.hex_decode("41g")    # => None

fn int_to_hex(n: Int) -> Str

str.int_to_hex(255)   # => ff

fn hex_to_int(s: Str) -> Option(Int)

"ff", "0xFF" and "-0x1a" all read; anything else is None.

str.hex_to_int("0xFF")    # => Some(255)
str.hex_to_int("-0x1a")   # => Some(-26)
str.hex_to_int("zz")      # => None

fn hex_digits(s: Str, i: Int) -> Option(Int)

The hex digits from i to the end of s, as a number; None if anything else is there.

str.hex_digits("#ff", 1)   # => Some(255)

fn url_encode(s: Str) -> Str

Percent-encoding as RFC 3986 has it: letters, digits and -_.~ stand, every other byte becomes %XX. form_encode is the variant HTML forms use, with a space as +.

str.url_encode("a b/c")   # => a%20b%2Fc

fn form_encode(s: Str) -> Str

str.form_encode("a b&c")   # => a+b%26c

fn url_decode(s: Str) -> Option(Str)

A % not followed by two hex digits is None.

str.url_decode("a%20b")   # => Some(a b)
str.url_decode("a%2")     # => None

fn form_decode(s: Str) -> Option(Str)

str.form_decode("a+b%26c")   # => Some(a b&c)

fn html_escape(s: Str) -> Str

The five characters HTML cannot show as they are. Safe for text and for attribute values in either kind of quote.

str.html_escape("<a href=\"x\">")   # => &lt;a href=&quot;x&quot;&gt;

fn html_unescape(s: Str) -> Str

The five named entities and the numeric ones, decimal and hex. An & that starts no entity is kept as it is, which is what browsers do.

str.html_unescape("&lt;b&gt; &amp; &#233; &#x41;")   # => <b> & é A

fn escape(s: Str) -> Str

C-style escapes: \n \r \t \\ \" and \xHH for any other control byte, so that a string can be shown on one line and read back exactly.

str.escape("a\tb\n")   # => a\tb\n

fn unescape(s: Str) -> Option(Str)

Reads escape's output back, plus \', \0 and \u{...}. A backslash that starts no escape is None.

str.unescape("a\\tb")     # => Some(a   b)
str.unescape("\\u{e9}")    # => Some(é)
str.unescape("\\q")        # => None

fn par_count(s: Str, p: Str) -> Int

count_all (every occurrence, overlapping ones included) over all the workers. The overlapping count is the one that splits: whether a non-overlapping scan takes a match depends on the matches before it, which another strand may hold. For a pattern that cannot overlap itself the two counts are the same number.

str.par_count("aaa", "aa")   # => 2

fn par_count_on(s: Str, p: Str, k: Int) -> Int

str.par_count_on("the cat, the hat", "the", 4)   # => 2

fn par_index_of(s: Str, p: Str) -> Int

index_of over all the workers: each range reports the first match that begins in it, and the smallest wins.

str.par_index_of("hello", "ll")   # => 2

fn par_index_on(s: Str, p: Str, k: Int) -> Int

str.par_index_on("hello", "lo", 3)   # => 3

fn par_lines(s: Str, f: (Str) -> 'a) -> List('a)

map(lines(s), f) over all the workers: the ranges are moved forward to the byte after a newline so that no line is cut, each strand maps its own lines, and the pieces are joined in order. The lines are exactly those of the prelude's lines, trailing empty one included.

str.par_lines("a\nbb\n", \l -> #l)   # => Cons(1, Cons(2, Cons(0, Nil)))

fn par_lines_on(s: Str, f: (Str) -> 'a, k: Int) -> List('a)

str.par_lines_on("a\nbb\nccc", \l -> #l, 2)   # => Cons(1, Cons(2, Cons(3, Nil)))