Strings — UTF-16, Runes, Interpolation, Immutability, StringBuffer

By the end of this lesson you will know exactly what a Dart string IS in memory (a row of numbers, not magic text), why an emoji can confuse .length, how $name interpolation really runs, why s += x in a loop can quietly wreck your program's speed, and the core toolbox of string methods, regular expressions, and classic string interview questions.

1. What a string really is (numbers in disguise)

Think of an old telegraph office. Nobody sends the LETTER "A" down the wire — they send a numeric code for it, and the receiving machine looks the code up in a table to print "A" on paper. A computer's memory works the same way: it can only store numbers, so every character you type gets converted to a number first. The lookup table that assigns numbers to characters is called a character encoding.

The oldest common table is ASCII: it assigns numbers 0–127 to English letters, digits and punctuation ('A' → 65, 'a' → 97, '0' → 48). ASCII only covers English, so the world agreed on a bigger table called Unicode, which assigns a unique number (called a code point) to over a million characters — every script on Earth, plus emoji.

A Dart String behaves as a sequence of UTF-16 code units — each one a 16-bit (0–65,535) number, using the UTF-16 encoding of Unicode. (Behaves as: the Dart VM may keep plain Latin-1 text in a more compact one-byte layout internally, but everything you can observe — length, codeUnitAt, codeUnits — is always in UTF-16 code units, so that is the model to use.) 'A'.codeUnitAt(0) gives you the raw number 65, and .codeUnits gives you the whole list of numbers behind the text.

The same idea as runnable Dart (the comment-free lines print what is shown after // =>):

Input size → what's feasible: a string of n ASCII letters is n numbers, so reading one code unit is O(1) and listing them all is O(n) — fine for n = 106 characters.

A String is not an array of "characters" the way beginners picture it — it's an array of 16-bit numbers. Most of the time that distinction doesn't matter (plain English text), but the moment you touch emoji, accented letters, or non-Latin scripts, it matters a lot (see section 2).
A string is a sequence of numbers (UTF-16 code units) plus a table that says how to draw each number as a glyph on screen. codeUnitAt(i) and codeUnits let you see the raw numbers directly.

2. Code units vs runes vs grapheme clusters

Picture a two-page fold-out photo in a magazine. If you count "pages", it's 2. But if you count "pictures", it's 1 — the two pages together form a single image. UTF-16 code units are like pages; Unicode code points (runes) are like the picture they combine to form.

Some Unicode code points (like 😀, code point U+1F600) are too big to fit in one 16-bit code unit, so UTF-16 splits them into a pair of two code units called a surrogate pair. That means '😀'.length (which counts UTF-16 code units) is 2, while '😀'.runes.length (which counts actual Unicode code points) is 1. Dart's .runes property gives you the real code points.

In code: the emoji's length, runes, and the exact UTF-16 formula that splits code point U+1F600 into the pair 55357 + 56832 (subtract 0x10000, then 10 high bits + 10 low bits):

There's a third, even higher level: a grapheme cluster — what a human actually perceives as "one character" on screen. A family emoji like 👨‍👩‍👧‍👦 or a letter with a combining accent mark (e.g. e + a combining acute accent forming "é") can be MULTIPLE runes that a person still reads as one glyph. Neither .length nor .runes.length counts grapheme clusters correctly — for that you need the separate characters package (import 'package:characters/characters.dart', .characters.length), which is not part of the core library, so it isn't exercised in this lesson's verification file, but it is the right tool whenever you must count, reverse, or slice text the way a human sees it (e.g. cursor movement in a text editor).

Both cases in Dart — a composed vs decomposed "é", and a family emoji that is 4 emoji + 3 joiners (7 runes, 11 code units) yet one glyph on screen:

Input size → what's feasible: length is O(1) and runes.length is O(n) (it must decode the whole string); both are fine for any realistic text (n ≤ 106); only counting grapheme clusters needs the extra characters package.

Never assume .length equals "number of letters a human would count". For plain ASCII text they agree; for emoji, accents, and many non-English scripts they do not.
Three ways to measure "how long" a string is, and they can all disagree: code units (.length, .codeUnits) — raw UTF-16 storage; runes (.runes.length) — real Unicode code points; grapheme clusters (characters package) — what a human perceives as one character.

3. Literals & interpolation

Interpolation is like a mail-merge letter: the template has blanks ($name), and before printing, the computer fills each blank in with the current value, converting it to text first if it isn't text already.

Dart string literals: single quotes 'hi' and double quotes "hi" are identical (pick one style and stay consistent). Escapes use backslash: '\n' (newline), '\t' (tab), '\'' (an escaped quote), '\u{1F600}' (a Unicode code point by number). A raw string r'C:\new\folder' turns escapes OFF — backslashes are just backslashes, handy for regular expressions and file paths. A multi-line string uses triple quotes: '''line one\nline two''' can also just contain real newlines. Two adjacent string literals with nothing but whitespace/newline between them are automatically concatenated: 'foo' 'bar' is the single string 'foobar'.

Interpolation: $identifier inserts a variable's value; ${expression} inserts the result of any expression. Dart evaluates the expression, calls .toString() on the result (unless it's already a String), and splices the text in — you never need to write + x.toString() + by hand.

Literals and interpolation, verified line by line:

Input size → what's feasible: interpolation builds one new string of length n in O(n); a literal or a few $x slots is instant, and even n = 106 characters is fine once (the trouble starts only when you rebuild in a loop — section 5).

$name.length interpolates ONLY $name, then appends the literal text .length — because Dart's simple-identifier form ($name) stops at the identifier boundary. To interpolate an expression like name.length, you must use braces: ${name.length}.

4. Immutability — why += makes a new string

Imagine a stone tablet with "cat" carved into it. You cannot un-carve a letter or add more to the same tablet — if you want "catdog", a stonemason has to carve a BRAND NEW tablet with all five letters, and the old "cat" tablet just sits there, unchanged, until nobody needs it anymore.

Every String in Dart is immutable: once created, its contents can never change. s += 'dog' (short for s = s + 'dog') does not edit the existing string in place — it builds a whole new string object containing the old characters plus the new ones, then makes the variable s point at that new object. Any other variable still pointing at the old string (like old below) keeps seeing the original, untouched value.

The same story in code. Strings are covered more deeply in D22 · Mutable vs Immutable (what "immutable" means for every Dart type):

Input size → what's feasible: one += on a string of n characters costs O(n) (it copies); once or twice is fine for n = 106, but inside a loop it multiplies — see section 5.

Because strings are immutable, Dart can safely share the same string object between many variables without fear that one owner will silently corrupt it for the others — this is one reason immutable data is popular in concurrent and functional-style code.

5. Building strings fast: loop += vs StringBuffer

Copying the tablet every single time you add one letter is fine for a 3-letter word, but imagine re-carving an entire 10,000-letter tablet from scratch every time you add ONE more letter, ten thousand times over. That's what s += x inside a loop does. A StringBuffer is a scratch notepad — you jot each new bit down at the end and only carve the final stone tablet once, at the very end (.toString()).

Each += inside a loop allocates a new string and copies every existing character into it, then adds the new piece. If the string grows to length i on iteration i, that iteration costs about i character-copies. Summed over n iterations that's 1 + 2 + ... + n = n(n+1)/2, which is O(n²) total work. StringBuffer.write() instead collects the pieces in an internal buffer and builds the final string only once, so each write is (amortized) O(1) — the whole loop is O(n). Honest hedge: this is the cost model on the Dart VM (every += allocates a new string). The exact internals are implementation details, and when Dart is compiled to JavaScript the browser engine may optimise += with internal tricks — but StringBuffer / join is the portable way that is fast everywhere.

The maths from the paragraph above as code: the loop's copy count equals n(n+1)/2 (Σ i = 1 + 2 + … + n) for n = 10, 100, 1000, while StringBuffer does n writes:

Input size → what's feasible: n ≤ 103 → += is fine (≈ 5·105 copies); n = 105 → 5·109 copies, too slow, use StringBuffer or join. Every other builder (padLeft, replaceAll, *, …) is catalogued in D24 · Every String method, including its += vs StringBuffer section.

This is a real, common performance bug: building a big string with for (...) { result += piece; } looks innocent but becomes painfully slow as piece count grows into the thousands. Always prefer StringBuffer (or List<String>.join()) when assembling a string across many steps.
String s = ''; for (...) s += x; → O(n²). final sb = StringBuffer(); for (...) sb.write(x); sb.toString(); → O(n).

6. Core methods toolbox

A single string, 'banana', run through the everyday toolbox:

The whole toolbox as one runnable function (every line's output is shown):

Input size → what's feasible: each of these methods is O(n) or better on a string of n code units, so n ≤ 106 is fine; the complete String API (every method, with signatures) is in D24.

substring(start, end) excludes end — just like array slicing. 'banana'.substring(1, 3) takes indices 1 and 2 only, giving 'an', NOT 'ana'.

More essentials: s == t compares strings by value (are all the code units equal?), never by identity — two separately-built strings with the same letters are ==. s.compareTo(t) compares lexicographically by code unit (like a dictionary, but using the numeric code, not "alphabet position") — so 'Z'.compareTo('a') is negative, because uppercase letters (65–90) all have smaller codes than lowercase letters (97–122): 'Z' < 'a'. s.isEmpty / s.isNotEmpty check length 0 directly (clearer than s.length == 0). s.trim() removes leading/trailing whitespace. s.replaceAll(a, b) replaces every occurrence. s.padLeft(width, char) pads on the left until the string reaches width. s.split(sep) cuts into a List<String>; list.join(sep) glues a list back into one string.

Equality vs identity vs ordering, with the sort that surprises beginners:

7. RegExp basics

A regular expression is a search pattern written as a mini-language, like telling a librarian "find me any run of one-or-more digits" instead of spelling out every possible number.

RegExp(r'\d+') compiles a pattern matching one-or-more digits (note the raw string r'...' so \d isn't mangled by Dart's own escape rules). pattern.hasMatch(text) returns true/false. pattern.allMatches(text) returns every match found, each with .group(0) for the matched text. text.replaceAllMapped(pattern, (m) => ...) lets you compute a custom replacement for every match, using the match object m.

Input size → what's feasible: one regex scan is O(n) for typical patterns, fine for n ≤ 106; nested quantifiers like (a+)+ can backtrack exponentially, so keep patterns simple on user input.

8. Converting between strings, ints and chars

int.parse('42') turns text into a number (throws FormatException on invalid text — see the null-safety/errors lessons for handling that safely). 42.toString() turns a number back into text; interpolation does this for you automatically. String.fromCharCode(65) builds a one-character string from a single numeric code unit — the inverse of codeUnitAt.

Input size → what's feasible: int.parse reads the digits once, O(d) for d digits; a native Dart int holds up to 9.2·1018 (19 digits), so longer digit strings need BigInt (on the web the safe range is only 253).

double.tryParse / int.tryParse return null instead of throwing when the text isn't a valid number — usually the safer choice for user input, once you've studied null safety (lesson D05).

Quiz

Interview questions

Cheat sheet

NeedUseNotes
Raw numeric code of a chars.codeUnitAt(i)UTF-16 code unit, not "the letter"
All code unitss.codeUnitsList<int>
Real Unicode code pointss.runesuse for emoji-safe length/reverse
Human-perceived characterss.charactersneeds package:characters
Insert a value into text'$x' / '${expr}'calls toString()
No-escape literalr'...'backslashes stay literal — great for regex
Build across many stepsStringBufferO(n) vs O(n²) for += in a loop
Slice (end excluded)s.substring(a, b)
Value equalitys == tnever identity for content comparison
Dictionary orderings.compareTo(t)by code unit — uppercase < lowercase
Pattern searchRegExp(r'...')hasMatch, allMatches, replaceAllMapped
Text ⇄ numberint.parse / toString()tryParse for safe parsing