Code Generation & Performance
By the end of this lesson you will be able to explain what build_runner actually does and read/predict a generated *.g.dart file, write a tiny code generator of your own, and reason about a Dart program's performance the right way: measuring before guessing, recognising when Big-O dominates the outcome, avoiding unnecessary allocations, and caching/memoizing correctly — all with appropriately hedged claims, since exact numbers always depend on your machine and workload.
1. Why code generation exists
In real Dart/Flutter apps, four kinds of boilerplate show up constantly: turning a class to/from JSON (for talking to a server), giving a class real value equality (==/hashCode, so two objects with the same fields compare equal — see D13), building "immutable copy with these fields changed" methods (copyWith), and wiring up type-safe routes/arguments for navigation. Writing all four by hand for every model class is exactly the kind of repetitive, mechanical work a generator should do instead — freeing you to review the OUTPUT rather than hand-write it, and guaranteeing every class follows the exact same pattern.
Here is exactly that boilerplate for a two-field class, written by hand. Run it and read the exact output:
The four kinds of boilerplate, written by hand
Each new field forces edits in 8 places; a generator makes all 8 from the one field declaration.
.dart files to disk, which are then compiled normally. At runtime there is no generator involved at all — just ordinary generated Dart code.2. The build_runner pipeline & json_serializable
build_runner) carries every source file past a row of inspectors (builders). Each inspector only cares about files with a specific stamp on them (an annotation like @JsonSerializable()). When an inspector recognises its stamp, it reads the file, writes a companion sheet of instructions, and staples it to the original via a part directive — the two files are now one library, compiled together.build_runner itself does no code generation — it is the orchestrator. The actual generation logic lives in packages like source_gen (a shared toolkit for writing Dart generators) plus a specific builder such as json_serializable (JSON) or freezed (immutable data classes, unions, copyWith). You run it with dart run build_runner build (or --watch to keep it running and regenerate automatically as you edit).
Now look at exactly what json_serializable's builder writes, line by line, for that same Person class:
The same idea as running code. runBuilders imitates the three decisions build_runner makes (which files are annotated, what output name was promised, write that file). It has no network or package dependencies, so you can run and change it:
build_runner's decisions in 20 lines
And this is the shape of the real pair of files. The part directive makes both halves ONE library, which is why generated code may read the private field _secret. Both halves are put in one file here, which gives the same visibility:
person.dart + person.g.dart, as one library
The real json_serializable output uses exactly these names (_$PersonFromJson, _$PersonToJson) and (json['age'] as num).toInt() for ints.
build_runner after adding or changing an @JsonSerializable() class, then being confused that _$PersonFromJson "doesn't exist". The generated file is not created automatically by the IDE or by flutter run — you (or a watch task) must explicitly generate it, and re-generate it every time the annotated class's shape changes.@JsonKey(name: 'server_name') and per class with things like fieldRename. The generator reads your class through Dart's own static analyzer — the same tool your IDE uses for autocomplete — so it always sees your class's real, current shape, never a guess.3. part files, freezed, macros (cancelled), and lighter alternatives
The part 'x.g.dart'; directive at the top of a source file is a promise: "somewhere, a file named exactly x.g.dart will be merged into THIS library, sharing its imports and privacy." It's why generated code can freely call private (_-prefixed) members of your class, and why the generated filename is never arbitrary — it must match the part directive exactly or the library fails to compile.
freezed builds on the same build_runner/source_gen pipeline to generate an entire immutable data class from a short "declaration-only" class you write: value equality, hashCode, a working toString(), a type-safe copyWith, and (its signature feature) tagged unions for modelling "one of several possible shapes" cleanly, all from one small annotated declaration. After macros were discontinued (below), the freezed 3.x line kept this same build_runner/source_gen foundation and added a "mixed mode" that lets a class mix hand-written members alongside generated ones — a real, current change worth knowing about if you read a freezed 2.x tutorial and then open a 3.x project, though the exact options are best checked against the version pinned in your own pubspec.yaml rather than memorised here.
Not every boilerplate problem needs a generator at all. Dart 3's own language features remove a lot of it directly: records ((String, int) — see D11) give you a lightweight, structurally-equal tuple with zero generated code; sealed classes with exhaustive switch patterns (also D11) replace a lot of what a hand-written "union type" used to need; extension types (a zero-cost compile-time wrapper around an existing type) can replace a small wrapper class that previously needed its own equality and forwarding boilerplate.
What freezed writes for you, spelled out by hand for a two-field class (so you can review generated code instead of trusting it):
Value equality + copyWith + toString, by hand
Macros were discontinued (January 2025), so there is no macro panel to run: the only supported generated-code path is the one shown in runBuilders above, which needs a real part file on disk.
And the three language features that often remove the need for a generator at all:
Record, sealed class, extension type (no generator, no build step)
4. Write a tiny generator yourself
source_gen or the analyzer package to understand the IDEA of code generation — at its core, a generator is just a function that takes a description of your data (field names and types) and returns a String of Dart source text. That's it. The "real" generators do the exact same thing; they just get their input by parsing your actual source files instead of a small hand-built FieldSpec list.In verify/d20.dart this lesson implements exactly that: a FieldSpec class describing one field (name + type), a generateJsonCode function that emits a toJson/fromJson pair as plain text, a generateEqualsHashCode function that emits ==/hashCode source from a field-name list, and a small generateTemplateFunction "mini template engine" that turns a {{placeholder}} string into real Dart function source. None of these COMPILE or RUN the text they produce — Dart has no runtime eval — so every claim about their output is checked as exact generated TEXT (the same way you'd review a real *.g.dart file by reading it), verified against output actually captured by running the functions. Note this hand-rolled version deliberately uses a simpler naming style (an extension, no _$ prefix) than the REAL json_serializable convention shown in animation 2 above (_$PersonFromJson/_$PersonToJson, no extension) — the idea is identical, only the exact output text differs, which is the point: a generator's job is "produce correct source text", however it happens to be styled. The full implementations and their exact expected output appear in the Medium- and Hard-tier questions below.
Run all three generators and read the generated text exactly as you would review a *.g.dart file (the functions themselves are the solutions to the Medium/Hard questions 14, 15 and 22 below):
Three tiny generators, their exact output
Input size → what is feasible: k ≤ 103 fields give ≈ 105 characters of output, built with a StringBuffer in O(output) time; += on a growing string would copy the output ≈ (105)²/(2·100) = 5·107 characters per run. The template function is O(m + v) for template length m ≤ 105 and v ≤ 103 variables.
generateTemplateFunction below deliberately throws an ArgumentError for an unknown placeholder rather than emitting broken source; this is tested directly in verify/d20.dart.5. Measure first: Stopwatch & DevTools
Dart's Stopwatch (dart:core) is the basic timing tool: Stopwatch()..start(), do work, .stop(), read .elapsedMicroseconds. A single Stopwatch reading of a cold function is close to meaningless on its own (section 7 explains exactly why) — real measurement means warming up first and averaging many iterations, which this lesson's Expert-tier "fix a bad benchmark" question below implements and tests. For a running Flutter or Dart VM app, DevTools gives you much richer, live tools: a CPU profiler (which functions actually consumed time, sampled while your app runs) and a Memory view (see D13) — both far more informative than eyeballing code and guessing.
The basics, with exact output. Only deterministic facts are printed with print; the microseconds go to stderr because they change every run:
Stopwatch basics
DevTools shows what you mark with dart:developer. A Timeline slice costs nothing when no tool is attached and never changes the result:
Marking a slice for the DevTools timeline
A benchmark done properly has three parts: warm-up calls whose timing is thrown away, many rounds each timing many calls, and the median of the rounds. Below it compares s += piece with StringBuffer. The exact ratio is printed to the console (stderr) instead of being asserted, because it differs per machine:
A proper micro-benchmark: warm-up, many runs, median
Measured on the author's machine (Apple M3 Pro, macOS, Dart 3.11.5; medians of 9 rounds after warm-up, run twice): 2000 pieces of 2 characters took about 225–265 µs with += and about 23–27 µs with StringBuffer, a ratio of about 9.4–9.8×. Your numbers will differ; the direction and the order of magnitude will not.
6. Big-O beats micro-optimisations
Concretely: checking whether a list of n numbers contains a duplicate with a nested loop (compare every pair) does roughly n·(n-1)/2 comparisons — O(n²). Doing the same check with a Set (insert each element; if an insert reports "already there", you found a duplicate) does exactly n set operations — O(n). Both are exercised and their EXACT operation counts checked (not guessed) in verify/d20.dart.
Counting steps instead of timing them
Σi=0n−1(n−1−i) = n(n−1)/2 is checked against the loop; "grows 10x" means n(n−1)/2 grows about 100x.
Input size → what is feasible: n ≤ 103 → the O(n²) all-pairs check is fine (5·105 steps); n = 105 → it needs 5·109 steps, too slow, use the O(n) Set check (105 steps).
7. JIT vs AOT
The Dart VM can run your code two different ways, and which one applies changes some performance intuitions (all of the following is hedged deliberately, since exact numbers are release- and platform-specific). In JIT mode (Just-In-Time — used for dart run and flutter run in debug mode, and for hot reload), Dart source is compiled to machine code WHILE the program runs, often starting from slower, unoptimised machine code and re-compiling "hot" (frequently executed) functions to more optimised machine code as it learns which ones matter — this is exactly the "warm-up" behaviour section 5 and the Expert-tier benchmarking question below are about. In AOT mode (Ahead-Of-Time — used for Flutter release builds and compiled Dart executables), all code is compiled to native machine code BEFORE the program ever runs, with no warm-up phase and no runtime recompilation, generally trading JIT's ability to re-optimise based on live behaviour for a predictable, fast start.
What the running program can tell you about its mode
Measured on the author's machine (same M3 Pro, Dart 3.11.5): one function summing 300 000 terms, timed 8 times in a row. Under dart run (JIT) the first two rounds took about 560–730 µs, then settled at about 195–210 µs (about 3× faster once optimised). The same code compiled with dart compile exe (AOT) took about 93–137 µs from the very first call and stayed flat. So warm-up is a JIT cost, and a single cold timing under dart run overstates the steady cost by about 3× here. Spikes (such as one round of 1288 µs in one run) are why you take a median.
8. Allocations in hot loops
new object created inside a hot loop is extra work for both the allocator (finding space) and, later, the garbage collector (reclaiming that space once it's unreachable — see D13). Reusing one mutable object instead of allocating a fresh one every iteration can remove that overhead entirely for the reused object.
Reuse vs allocate, and the ring buffer
1 object vs one per offset; the ring keeps ONE fixed array and overwrites the oldest value when full.
Input size → what is feasible: n ≤ 107 offsets → both loops are fine in time; the reuse version allocates 1 object, the immutable one allocates n (107 short-lived objects for the garbage collector).
String building is a very common accidental hot loop: repeatedly using += on a String inside a loop, because Dart strings are immutable — every += allocates an entirely new string and copies everything accumulated SO FAR into it, so total copying work across n concatenations is 1+2+...+n, which is O(n²). A StringBuffer (introduced in D04) collects fragments internally and builds the final string exactly once, doing O(n) total work — this equivalence (same output, different total work) is tested directly in verify/d20.dart, never asserted via timing.
Counting the copying: += vs StringBuffer
n(n+1)/2 characters copied by += versus n writes plus one final copy.
Input size → what is feasible: n ≤ 103 pieces → either way; n = 2·105 pieces of 10 characters → += copies ≈ 2·1011 characters (too slow), StringBuffer copies ≈ 2·106.
Two more allocation-related tools worth knowing: typed data (Uint8List, Float64List, etc., covered in D16) stores numbers packed contiguously in a fixed-size buffer instead of as individually-boxed objects in a general List — useful for large numeric buffers (audio/image/binary data) where per-element object overhead would otherwise add up. And @pragma('vm:prefer-inline') is a VM-specific hint you CAN place above a small, extremely hot function to suggest inlining it at its call sites, removing call overhead — but this is a narrow, implementation-specific hint that the compiler is free to ignore, it can make code SLOWER if misapplied (larger inlined code hurts instruction-cache behaviour), and it should only ever follow real profiling evidence that a specific tiny function's call overhead actually matters, never be added speculatively.
Typed data: same answer, smaller storage
The inline hint compiles and changes nothing about the answer
Σ i² = n(n+1)(2n+1)/6 checked for n = 999.
Measured (M3 Pro, Dart 3.11.5, 105 elements, median of 9): summing a List<int> took about 54–56 µs and an Int32List about 40–43 µs, a ratio of about 1.3–1.4×. This was with an indexed loop (for (var i = 0; i < n; i++) s += xs[i]); on the same machine a for (final v in typedList) loop was slower, not faster, than the same loop over a List<int>, so loop style can matter as much as the type. The big win of typed data is memory and I/O layout, not loop speed.
9. const, collections, dynamic, and isolates
A const constructor call (e.g. const SizedBox(height: 8)) is canonicalized at compile time (see D13's identity lesson) — every textually-identical const expression reuses the exact same object, so a Flutter widget tree that reuses const widgets can skip re-creating (and, for widgets, potentially re-rendering) them. The exact rebuild-avoidance behaviour is a Flutter framework detail, hedged here deliberately — the language-level guarantee is only about object identity/canonicalization, not about frame timing.
const gives one shared object
Choosing the right collection (D08) for your access pattern matters more than almost any micro-optimisation: a List gives O(1) index access but O(n) "is this value present" checks; a Set/Map gives average O(1) membership/lookup by hashing, at the cost of not preserving a meaningful numeric index; a LinkedHashMap/LinkedHashSet keeps insertion order alongside hash-based lookup. Using .contains() on a List inside a loop (turning an intended O(n) algorithm into O(n²)) is one of the single most common real-world performance bugs — see the Hard-tier dedupe question below for exactly this fix.
Counting == calls: List.contains vs Set.contains
Worst-case lookup of the last element: n comparisons for a list, 1 for a hash set.
Measured (5000 lookups in a 5000-element collection, median): List.contains about 6900–7200 µs versus Set.contains about 13–28 µs, a ratio of about 250–540×. As a queue, removeAt(0) on a 2·104-element list took about 110 ms versus about 0.22–0.26 ms with ListQueue.removeFirst, about 420–530×: never use removeAt(0) as a queue.
A chain like xs.where(...).map(...).toList() uses lazy iterables — each stage doesn't build an intermediate list, it produces values on demand as the next stage asks for them. This avoids allocating throwaway intermediate lists, which is good, but chaining many such stages does add some per-element call/wrapper overhead compared to one hand-written loop doing the same filtering and mapping directly; for a hot loop processing millions of elements this is worth knowing, but for everyday code the clarity of .where().map() is usually worth far more than any difference here, which is typically small and always workload/version dependent.
Lazy chains: counting what actually runs
Measured (105 elements, filter even then double then sum): a hand-written loop took about 70–95 µs, the where().map().fold() chain about 600–715 µs, so the chain was about 7–9× slower on this workload: clarity usually wins, but not inside a hot loop over millions of elements.
Using dynamic turns off Dart's compile-time type checking for that value, and a call through a dynamic-typed reference generally cannot use the same fast, statically-resolved dispatch a concretely-typed call can — it's typically somewhat slower, though exactly how much depends on the VM/compiler version and is not something to hard-code a number for. The much bigger cost of overusing dynamic is usually correctness, not speed: you lose the compiler's help catching a typo'd method name or a wrong-type argument until runtime.
dynamic: the typo is found at run time
Measured (indexing a 105-element List<int>): through a dynamic variable 1.0× (no measurable difference, about 61–70 µs both ways), because the receiver is always the same type and the VM specialises the call. The cost of dynamic is the lost compile-time check, not a guaranteed slowdown.
For genuinely CPU-heavy work (not I/O — I/O already doesn't block Dart's event loop, see D14), moving it into a separate isolate (D15) is the real tool: it lets the main isolate's event loop stay responsive to input/UI/other work while the heavy computation runs elsewhere. This doesn't make the underlying algorithm any faster by itself (same Big-O, same total work) — it changes WHERE that work happens so it doesn't block anything else, and can add real parallelism if you split the work across several isolates on a multi-core machine.
Isolate.run: same work, same answer, different place
10. Flutter-specific notes (brief, hedged)
Flutter's own performance story deserves a full lesson of its own; briefly: a rebuild (Flutter re-running a widget's build() method) is generally cheap by itself, but doing expensive work directly inside build() multiplies that cost by every rebuild. A const widget constructor lets Flutter skip rebuilding that specific subtree when its parent rebuilds, since the framework can tell nothing about it changed. Jank means a visibly dropped/late frame — commonly described as a frame that misses its budget (roughly 16ms for a smooth 60Hz display, less on higher refresh-rate screens), causing visible stutter; the exact budget depends on the display's actual refresh rate and is best treated as "keep each frame's work small," not a hard-coded constant to test against in this lesson's non-UI verify file.
The frame budget is just arithmetic
11. Caching, memoization & precomputation
f(4)? It was 16 — here, don't redo the work." It only works safely for a pure function — one that always returns the same output for the same input and has no other effects — because the sticky note is trusted blindly on every repeat call.The classic teaching example is a naive recursive Fibonacci, which recomputes the same smaller sub-problems an exploding number of times; memoizing it (caching each n's result the first time it's computed) turns an exponential blow-up into roughly one computation per distinct n:
Naive vs memoized: exact call counts and the closed form
calls(n) = 1 + calls(n−1) + calls(n−2) solves to 2·fib(n+1) − 1; memoized needs 2n−1 calls for n ≥ 2.
Input size → what is feasible: n ≤ 35 → naive fib is ≈ 3·107 calls (fits); n = 50 → ≈ 4·1010 calls (too slow); the memoized version handles n ≤ 90 (the result fits 64 bits) with recursion depth ≤ 91. Measured (fib(25), median): naive about 310–556 µs, memoized about 0.9–2.2 µs, about 240–390×.
Precomputation is memoization's cousin: instead of caching answers lazily as they're first asked for, you compute a useful structure UP FRONT, once, so every later query is cheap. The prefix-sum technique in this lesson's Medium-tier "optimize this loop" question below is a clean example: computing a running-total array once in O(n) turns every subsequent range-sum query into O(1), instead of re-summing the range from scratch every time.
Prefix sums: the formula as code, with the 1-indexed → 0-indexed shift
Math form: sum(l..r) = P[r] − P[l−1] (1-indexed). Dart form: prefix[hi+1] − prefix[lo] with lo = l−1, hi = r−1.
Input size → what is feasible: n, q ≤ 2·105 → re-summing is n·q = 4·1010 steps (too slow); prefix sums are n + q = 4·105.
A general-purpose cache extends memoization with real-world bounds: it can't grow forever (an LRU — Least Recently Used — policy evicts the entry that hasn't been touched in the longest time once the cache is full) and entries can go stale (a TTL — Time To Live — expires an entry after a set duration even if it's never evicted for space). The Hard-tier "cache with TTL + LRU" question below implements and tests both policies together, using an injectable fake clock so the TTL behaviour is fully deterministic and never depends on a real sleep() or wall-clock timing.
TTL + LRU with a fake clock
What goes wrong when you memoize an impure function
Quiz
Interview questions
Cheat sheet
| Term | Meaning |
|---|---|
build_runner | Orchestrates the build: scans source files, dispatches matching ones to registered builders (e.g. json_serializable, freezed) |
part 'x.g.dart'; | Promises a generated file will be merged into this library; names must match exactly |
| macros (Dart) | Experimental compiler-integrated codegen feature — officially discontinued January 2025; not used in current Dart |
| Big-O | Describes how work GROWS with input size; dominates over constant-factor micro-optimisations for large enough input |
| JIT | Compiles to machine code while running, with a warm-up phase; used in debug, dart run and hot reload |
| AOT | Compiles to machine code before running, no warm-up; used in Flutter release builds |
| Hot loop | A loop that runs enough times that per-iteration cost (e.g. allocations) matters |
StringBuffer | Builds a string in O(n) total work; avoids the O(n²) cost of repeated += |
| Memoization | Caching a pure function's results by input, lazily, on first call |
| Precomputation | Building a useful structure up front so later queries are cheap (e.g. prefix sums) |
| LRU / TTL cache | Evicts the least-recently-used entry when full / expires entries after a fixed duration |
| Performance budget | A deterministic, operation-count-based regression test — never a wall-clock assertion |