Strings — UTF-16, Runes, Interpolation, Immutability, StringBuffer
By the end of this lesson you will know exactly what a Dart string IS in memory (a row of numbers, not magic text), why an emoji can confuse .length, how $name interpolation really runs, why s += x in a loop can quietly wreck your program's speed, and the core toolbox of string methods, regular expressions, and classic string interview questions.
1. What a string really is (numbers in disguise)
The oldest common table is ASCII: it assigns numbers 0–127 to English letters, digits and punctuation ('A' → 65, 'a' → 97, '0' → 48). ASCII only covers English, so the world agreed on a bigger table called Unicode, which assigns a unique number (called a code point) to over a million characters — every script on Earth, plus emoji.
A Dart String behaves as a sequence of UTF-16 code units — each one a 16-bit (0–65,535) number, using the UTF-16 encoding of Unicode. (Behaves as: the Dart VM may keep plain Latin-1 text in a more compact one-byte layout internally, but everything you can observe — length, codeUnitAt, codeUnits — is always in UTF-16 code units, so that is the model to use.) 'A'.codeUnitAt(0) gives you the raw number 65, and .codeUnits gives you the whole list of numbers behind the text.
The same idea as runnable Dart (the comment-free lines print what is shown after // =>):
Input size → what's feasible: a string of n ASCII letters is n numbers, so reading one code unit is O(1) and listing them all is O(n) — fine for n = 106 characters.
String is not an array of "characters" the way beginners picture it — it's an array of 16-bit numbers. Most of the time that distinction doesn't matter (plain English text), but the moment you touch emoji, accented letters, or non-Latin scripts, it matters a lot (see section 2).codeUnitAt(i) and codeUnits let you see the raw numbers directly.2. Code units vs runes vs grapheme clusters
Some Unicode code points (like 😀, code point U+1F600) are too big to fit in one 16-bit code unit, so UTF-16 splits them into a pair of two code units called a surrogate pair. That means '😀'.length (which counts UTF-16 code units) is 2, while '😀'.runes.length (which counts actual Unicode code points) is 1. Dart's .runes property gives you the real code points.
In code: the emoji's length, runes, and the exact UTF-16 formula that splits code point U+1F600 into the pair 55357 + 56832 (subtract 0x10000, then 10 high bits + 10 low bits):
There's a third, even higher level: a grapheme cluster — what a human actually perceives as "one character" on screen. A family emoji like 👨👩👧👦 or a letter with a combining accent mark (e.g. e + a combining acute accent forming "é") can be MULTIPLE runes that a person still reads as one glyph. Neither .length nor .runes.length counts grapheme clusters correctly — for that you need the separate characters package (import 'package:characters/characters.dart', .characters.length), which is not part of the core library, so it isn't exercised in this lesson's verification file, but it is the right tool whenever you must count, reverse, or slice text the way a human sees it (e.g. cursor movement in a text editor).
Both cases in Dart — a composed vs decomposed "é", and a family emoji that is 4 emoji + 3 joiners (7 runes, 11 code units) yet one glyph on screen:
Input size → what's feasible: length is O(1) and runes.length is O(n) (it must decode the whole string); both are fine for any realistic text (n ≤ 106); only counting grapheme clusters needs the extra characters package.
.length equals "number of letters a human would count". For plain ASCII text they agree; for emoji, accents, and many non-English scripts they do not..length, .codeUnits) — raw UTF-16 storage; runes (.runes.length) — real Unicode code points; grapheme clusters (characters package) — what a human perceives as one character.3. Literals & interpolation
$name), and before printing, the computer fills each blank in with the current value, converting it to text first if it isn't text already.Dart string literals: single quotes 'hi' and double quotes "hi" are identical (pick one style and stay consistent). Escapes use backslash: '\n' (newline), '\t' (tab), '\'' (an escaped quote), '\u{1F600}' (a Unicode code point by number). A raw string r'C:\new\folder' turns escapes OFF — backslashes are just backslashes, handy for regular expressions and file paths. A multi-line string uses triple quotes: '''line one\nline two''' can also just contain real newlines. Two adjacent string literals with nothing but whitespace/newline between them are automatically concatenated: 'foo' 'bar' is the single string 'foobar'.
Interpolation: $identifier inserts a variable's value; ${expression} inserts the result of any expression. Dart evaluates the expression, calls .toString() on the result (unless it's already a String), and splices the text in — you never need to write + x.toString() + by hand.
Literals and interpolation, verified line by line:
Input size → what's feasible: interpolation builds one new string of length n in O(n); a literal or a few $x slots is instant, and even n = 106 characters is fine once (the trouble starts only when you rebuild in a loop — section 5).
$name.length interpolates ONLY $name, then appends the literal text .length — because Dart's simple-identifier form ($name) stops at the identifier boundary. To interpolate an expression like name.length, you must use braces: ${name.length}.4. Immutability — why += makes a new string
Every String in Dart is immutable: once created, its contents can never change. s += 'dog' (short for s = s + 'dog') does not edit the existing string in place — it builds a whole new string object containing the old characters plus the new ones, then makes the variable s point at that new object. Any other variable still pointing at the old string (like old below) keeps seeing the original, untouched value.
The same story in code. Strings are covered more deeply in D22 · Mutable vs Immutable (what "immutable" means for every Dart type):
Input size → what's feasible: one += on a string of n characters costs O(n) (it copies); once or twice is fine for n = 106, but inside a loop it multiplies — see section 5.
5. Building strings fast: loop += vs StringBuffer
s += x inside a loop does. A StringBuffer is a scratch notepad — you jot each new bit down at the end and only carve the final stone tablet once, at the very end (.toString()).Each += inside a loop allocates a new string and copies every existing character into it, then adds the new piece. If the string grows to length i on iteration i, that iteration costs about i character-copies. Summed over n iterations that's 1 + 2 + ... + n = n(n+1)/2, which is O(n²) total work. StringBuffer.write() instead collects the pieces in an internal buffer and builds the final string only once, so each write is (amortized) O(1) — the whole loop is O(n). Honest hedge: this is the cost model on the Dart VM (every += allocates a new string). The exact internals are implementation details, and when Dart is compiled to JavaScript the browser engine may optimise += with internal tricks — but StringBuffer / join is the portable way that is fast everywhere.
The maths from the paragraph above as code: the loop's copy count equals n(n+1)/2 (Σ i = 1 + 2 + … + n) for n = 10, 100, 1000, while StringBuffer does n writes:
Input size → what's feasible: n ≤ 103 → += is fine (≈ 5·105 copies); n = 105 → 5·109 copies, too slow, use StringBuffer or join. Every other builder (padLeft, replaceAll, *, …) is catalogued in D24 · Every String method, including its += vs StringBuffer section.
for (...) { result += piece; } looks innocent but becomes painfully slow as piece count grows into the thousands. Always prefer StringBuffer (or List<String>.join()) when assembling a string across many steps.String s = ''; for (...) s += x; → O(n²). final sb = StringBuffer(); for (...) sb.write(x); sb.toString(); → O(n).6. Core methods toolbox
A single string, 'banana', run through the everyday toolbox:
The whole toolbox as one runnable function (every line's output is shown):
Input size → what's feasible: each of these methods is O(n) or better on a string of n code units, so n ≤ 106 is fine; the complete String API (every method, with signatures) is in D24.
substring(start, end) excludes end — just like array slicing. 'banana'.substring(1, 3) takes indices 1 and 2 only, giving 'an', NOT 'ana'.More essentials: s == t compares strings by value (are all the code units equal?), never by identity — two separately-built strings with the same letters are ==. s.compareTo(t) compares lexicographically by code unit (like a dictionary, but using the numeric code, not "alphabet position") — so 'Z'.compareTo('a') is negative, because uppercase letters (65–90) all have smaller codes than lowercase letters (97–122): 'Z' < 'a'. s.isEmpty / s.isNotEmpty check length 0 directly (clearer than s.length == 0). s.trim() removes leading/trailing whitespace. s.replaceAll(a, b) replaces every occurrence. s.padLeft(width, char) pads on the left until the string reaches width. s.split(sep) cuts into a List<String>; list.join(sep) glues a list back into one string.
Equality vs identity vs ordering, with the sort that surprises beginners:
7. RegExp basics
RegExp(r'\d+') compiles a pattern matching one-or-more digits (note the raw string r'...' so \d isn't mangled by Dart's own escape rules). pattern.hasMatch(text) returns true/false. pattern.allMatches(text) returns every match found, each with .group(0) for the matched text. text.replaceAllMapped(pattern, (m) => ...) lets you compute a custom replacement for every match, using the match object m.
Input size → what's feasible: one regex scan is O(n) for typical patterns, fine for n ≤ 106; nested quantifiers like (a+)+ can backtrack exponentially, so keep patterns simple on user input.
8. Converting between strings, ints and chars
int.parse('42') turns text into a number (throws FormatException on invalid text — see the null-safety/errors lessons for handling that safely). 42.toString() turns a number back into text; interpolation does this for you automatically. String.fromCharCode(65) builds a one-character string from a single numeric code unit — the inverse of codeUnitAt.
Input size → what's feasible: int.parse reads the digits once, O(d) for d digits; a native Dart int holds up to 9.2·1018 (19 digits), so longer digit strings need BigInt (on the web the safe range is only 253).
double.tryParse / int.tryParse return null instead of throwing when the text isn't a valid number — usually the safer choice for user input, once you've studied null safety (lesson D05).Quiz
Interview questions
Cheat sheet
| Need | Use | Notes |
|---|---|---|
| Raw numeric code of a char | s.codeUnitAt(i) | UTF-16 code unit, not "the letter" |
| All code units | s.codeUnits | List<int> |
| Real Unicode code points | s.runes | use for emoji-safe length/reverse |
| Human-perceived characters | s.characters | needs package:characters |
| Insert a value into text | '$x' / '${expr}' | calls toString() |
| No-escape literal | r'...' | backslashes stay literal — great for regex |
| Build across many steps | StringBuffer | O(n) vs O(n²) for += in a loop |
| Slice (end excluded) | s.substring(a, b) | |
| Value equality | s == t | never identity for content comparison |
| Dictionary ordering | s.compareTo(t) | by code unit — uppercase < lowercase |
| Pattern search | RegExp(r'...') | hasMatch, allMatches, replaceAllMapped |
| Text ⇄ number | int.parse / toString() | tryParse for safe parsing |