Code Generation & Performance

By the end of this lesson you will be able to explain what build_runner actually does and read/predict a generated *.g.dart file, write a tiny code generator of your own, and reason about a Dart program's performance the right way: measuring before guessing, recognising when Big-O dominates the outcome, avoiding unnecessary allocations, and caching/memoizing correctly — all with appropriately hedged claims, since exact numbers always depend on your machine and workload.

1. Why code generation exists

Imagine a form-filling clerk who, for every new employee, must copy the SAME ten fields from one paper form onto five other paper forms — a JSON form, an equality-check form, a "make me immutable" form, a routing form. The clerk never makes a creative decision; every field maps the same mechanical way every time. That is boilerplate: correct, necessary, utterly repetitive code. A human clerk doing this all day will eventually mistype a field. A code GENERATOR — a small program that reads your class's shape and writes that repetitive code for you — never does.

In real Dart/Flutter apps, four kinds of boilerplate show up constantly: turning a class to/from JSON (for talking to a server), giving a class real value equality (==/hashCode, so two objects with the same fields compare equal — see D13), building "immutable copy with these fields changed" methods (copyWith), and wiring up type-safe routes/arguments for navigation. Writing all four by hand for every model class is exactly the kind of repetitive, mechanical work a generator should do instead — freeing you to review the OUTPUT rather than hand-write it, and guaranteeing every class follows the exact same pattern.

Here is exactly that boilerplate for a two-field class, written by hand. Run it and read the exact output:

The four kinds of boilerplate, written by hand

Each new field forces edits in 8 places; a generator makes all 8 from the one field declaration.

Code generation in Dart never runs "live" inside your app. It's a build-time step: a separate program reads your annotated source files and writes plain .dart files to disk, which are then compiled normally. At runtime there is no generator involved at all — just ordinary generated Dart code.

2. The build_runner pipeline & json_serializable

Think of a small factory line. A conveyor belt (build_runner) carries every source file past a row of inspectors (builders). Each inspector only cares about files with a specific stamp on them (an annotation like @JsonSerializable()). When an inspector recognises its stamp, it reads the file, writes a companion sheet of instructions, and staples it to the original via a part directive — the two files are now one library, compiled together.

build_runner itself does no code generation — it is the orchestrator. The actual generation logic lives in packages like source_gen (a shared toolkit for writing Dart generators) plus a specific builder such as json_serializable (JSON) or freezed (immutable data classes, unions, copyWith). You run it with dart run build_runner build (or --watch to keep it running and regenerate automatically as you edit).

Now look at exactly what json_serializable's builder writes, line by line, for that same Person class:

The same idea as running code. runBuilders imitates the three decisions build_runner makes (which files are annotated, what output name was promised, write that file). It has no network or package dependencies, so you can run and change it:

build_runner's decisions in 20 lines

And this is the shape of the real pair of files. The part directive makes both halves ONE library, which is why generated code may read the private field _secret. Both halves are put in one file here, which gives the same visibility:

person.dart + person.g.dart, as one library

The real json_serializable output uses exactly these names (_$PersonFromJson, _$PersonToJson) and (json['age'] as num).toInt() for ints.

A very common early mistake: forgetting to actually RUN build_runner after adding or changing an @JsonSerializable() class, then being confused that _$PersonFromJson "doesn't exist". The generated file is not created automatically by the IDE or by flutter run — you (or a watch task) must explicitly generate it, and re-generate it every time the annotated class's shape changes.
json_serializable decides each field's JSON key using the field's own Dart name by default (as animation 2 shows), but this is configurable per field with @JsonKey(name: 'server_name') and per class with things like fieldRename. The generator reads your class through Dart's own static analyzer — the same tool your IDE uses for autocomplete — so it always sees your class's real, current shape, never a guess.

3. part files, freezed, macros (cancelled), and lighter alternatives

The part 'x.g.dart'; directive at the top of a source file is a promise: "somewhere, a file named exactly x.g.dart will be merged into THIS library, sharing its imports and privacy." It's why generated code can freely call private (_-prefixed) members of your class, and why the generated filename is never arbitrary — it must match the part directive exactly or the library fails to compile.

freezed builds on the same build_runner/source_gen pipeline to generate an entire immutable data class from a short "declaration-only" class you write: value equality, hashCode, a working toString(), a type-safe copyWith, and (its signature feature) tagged unions for modelling "one of several possible shapes" cleanly, all from one small annotated declaration. After macros were discontinued (below), the freezed 3.x line kept this same build_runner/source_gen foundation and added a "mixed mode" that lets a class mix hand-written members alongside generated ones — a real, current change worth knowing about if you read a freezed 2.x tutorial and then open a 3.x project, though the exact options are best checked against the version pinned in your own pubspec.yaml rather than memorised here.

A real, important fact to know for interviews and for reading Dart release notes: Dart's experimental macros feature — which was being developed specifically to let code generation happen without a separate build step, as compiler-integrated metaprogramming — was officially discontinued by the Dart team in January 2025. It never shipped as a stable feature. build_runner + source_gen remains the real, current way to do code generation in Dart as of this lesson; don't design new code around macros existing.

Not every boilerplate problem needs a generator at all. Dart 3's own language features remove a lot of it directly: records ((String, int) — see D11) give you a lightweight, structurally-equal tuple with zero generated code; sealed classes with exhaustive switch patterns (also D11) replace a lot of what a hand-written "union type" used to need; extension types (a zero-cost compile-time wrapper around an existing type) can replace a small wrapper class that previously needed its own equality and forwarding boilerplate.

What freezed writes for you, spelled out by hand for a two-field class (so you can review generated code instead of trusting it):

Value equality + copyWith + toString, by hand

Macros were discontinued (January 2025), so there is no macro panel to run: the only supported generated-code path is the one shown in runBuilders above, which needs a real part file on disk.

And the three language features that often remove the need for a generator at all:

Record, sealed class, extension type (no generator, no build step)

Reach for a code generator when the boilerplate is genuinely repetitive AND the shape is complex enough that hand-writing it invites mistakes (JSON (de)serialization of many fields, tagged unions). Reach for a plain Dart 3 language feature (record, sealed class, extension type) when a lighter built-in tool already solves the same problem with no build step at all.

4. Write a tiny generator yourself

You don't need source_gen or the analyzer package to understand the IDEA of code generation — at its core, a generator is just a function that takes a description of your data (field names and types) and returns a String of Dart source text. That's it. The "real" generators do the exact same thing; they just get their input by parsing your actual source files instead of a small hand-built FieldSpec list.

In verify/d20.dart this lesson implements exactly that: a FieldSpec class describing one field (name + type), a generateJsonCode function that emits a toJson/fromJson pair as plain text, a generateEqualsHashCode function that emits ==/hashCode source from a field-name list, and a small generateTemplateFunction "mini template engine" that turns a {{placeholder}} string into real Dart function source. None of these COMPILE or RUN the text they produce — Dart has no runtime eval — so every claim about their output is checked as exact generated TEXT (the same way you'd review a real *.g.dart file by reading it), verified against output actually captured by running the functions. Note this hand-rolled version deliberately uses a simpler naming style (an extension, no _$ prefix) than the REAL json_serializable convention shown in animation 2 above (_$PersonFromJson/_$PersonToJson, no extension) — the idea is identical, only the exact output text differs, which is the point: a generator's job is "produce correct source text", however it happens to be styled. The full implementations and their exact expected output appear in the Medium- and Hard-tier questions below.

Run all three generators and read the generated text exactly as you would review a *.g.dart file (the functions themselves are the solutions to the Medium/Hard questions 14, 15 and 22 below):

Three tiny generators, their exact output

Input size → what is feasible: k ≤ 103 fields give ≈ 105 characters of output, built with a StringBuffer in O(output) time; += on a growing string would copy the output ≈ (105)²/(2·100) = 5·107 characters per run. The template function is O(m + v) for template length m ≤ 105 and v ≤ 103 variables.

A generator that silently accepts a typo (e.g. a template placeholder that doesn't match any declared field) is worse than useless — it produces confidently-wrong code. generateTemplateFunction below deliberately throws an ArgumentError for an unknown placeholder rather than emitting broken source; this is tested directly in verify/d20.dart.

5. Measure first: Stopwatch & DevTools

A doctor doesn't prescribe treatment by guessing which organ "feels slow" — they run tests first. Performance work follows the identical rule: NEVER optimise code based on a hunch about what's slow. Measure, THEN optimise the part the measurement actually points at. Optimising the wrong 5% of a program can cost real engineering time for zero user-visible benefit.

Dart's Stopwatch (dart:core) is the basic timing tool: Stopwatch()..start(), do work, .stop(), read .elapsedMicroseconds. A single Stopwatch reading of a cold function is close to meaningless on its own (section 7 explains exactly why) — real measurement means warming up first and averaging many iterations, which this lesson's Expert-tier "fix a bad benchmark" question below implements and tests. For a running Flutter or Dart VM app, DevTools gives you much richer, live tools: a CPU profiler (which functions actually consumed time, sampled while your app runs) and a Memory view (see D13) — both far more informative than eyeballing code and guessing.

The basics, with exact output. Only deterministic facts are printed with print; the microseconds go to stderr because they change every run:

Stopwatch basics

DevTools shows what you mark with dart:developer. A Timeline slice costs nothing when no tool is attached and never changes the result:

Marking a slice for the DevTools timeline

A benchmark done properly has three parts: warm-up calls whose timing is thrown away, many rounds each timing many calls, and the median of the rounds. Below it compares s += piece with StringBuffer. The exact ratio is printed to the console (stderr) instead of being asserted, because it differs per machine:

A proper micro-benchmark: warm-up, many runs, median

Measured on the author's machine (Apple M3 Pro, macOS, Dart 3.11.5; medians of 9 rounds after warm-up, run twice): 2000 pieces of 2 characters took about 225–265 µs with += and about 23–27 µs with StringBuffer, a ratio of about 9.4–9.8×. Your numbers will differ; the direction and the order of magnitude will not.

"I think this is slow" is a hypothesis, not a finding. A Stopwatch reading, a DevTools CPU profile, or an operation-count test (section 11 / the Expert-tier performance-budget question below) is a finding. Optimise from findings.

6. Big-O beats micro-optimisations

Shaving 10% off the time each step of a walk takes is nice. Taking a route with 100× fewer steps is transformative. A micro-optimisation (a slightly faster line of code) usually buys a small, constant improvement. Choosing a better algorithm — one with a fundamentally smaller Big-O growth rate — can buy an improvement that gets LARGER, without limit, as your input grows. For any input big enough, the algorithmic win always eventually wins.

Concretely: checking whether a list of n numbers contains a duplicate with a nested loop (compare every pair) does roughly n·(n-1)/2 comparisons — O(n²). Doing the same check with a Set (insert each element; if an insert reports "already there", you found a duplicate) does exactly n set operations — O(n). Both are exercised and their EXACT operation counts checked (not guessed) in verify/d20.dart.

Counting steps instead of timing them

Σi=0n−1(n−1−i) = n(n−1)/2 is checked against the loop; "grows 10x" means n(n−1)/2 grows about 100x.

Input size → what is feasible: n ≤ 103 → the O(n²) all-pairs check is fine (5·105 steps); n = 105 → it needs 5·109 steps, too slow, use the O(n) Set check (105 steps).

A classic trap: micro-optimising the INSIDE of an O(n²) loop (e.g. making one comparison marginally cheaper) instead of noticing the loop itself is the wrong shape. A 2× faster comparison inside an O(n²) loop is still O(n²) — at n = 100,000 it is still doing on the order of 5 billion comparisons instead of 100,000. Fix the algorithm first; micro-optimise only what's left afterwards, and only once you've measured that it still matters.
Big-O describes how work GROWS as input grows, not an exact number of operations or a guaranteed wall-clock time (two different O(n) algorithms can still have very different constant-factor speeds). But when input size is even moderately large, the growth RATE almost always dominates every constant-factor concern — this is why "pick the right algorithm" beats "shave microseconds" as a general strategy.

7. JIT vs AOT

The Dart VM can run your code two different ways, and which one applies changes some performance intuitions (all of the following is hedged deliberately, since exact numbers are release- and platform-specific). In JIT mode (Just-In-Time — used for dart run and flutter run in debug mode, and for hot reload), Dart source is compiled to machine code WHILE the program runs, often starting from slower, unoptimised machine code and re-compiling "hot" (frequently executed) functions to more optimised machine code as it learns which ones matter — this is exactly the "warm-up" behaviour section 5 and the Expert-tier benchmarking question below are about. In AOT mode (Ahead-Of-Time — used for Flutter release builds and compiled Dart executables), all code is compiled to native machine code BEFORE the program ever runs, with no warm-up phase and no runtime recompilation, generally trading JIT's ability to re-optimise based on live behaviour for a predictable, fast start.

What the running program can tell you about its mode

Measured on the author's machine (same M3 Pro, Dart 3.11.5): one function summing 300 000 terms, timed 8 times in a row. Under dart run (JIT) the first two rounds took about 560–730 µs, then settled at about 195–210 µs (about 3× faster once optimised). The same code compiled with dart compile exe (AOT) took about 93–137 µs from the very first call and stayed flat. So warm-up is a JIT cost, and a single cold timing under dart run overstates the steady cost by about 3× here. Spikes (such as one round of 1288 µs in one run) are why you take a median.

Because of this, a debug/JIT build's performance is not a reliable stand-in for how a release/AOT build will behave — some code is relatively slower under JIT before warm-up, and the optimising compiler's decisions can differ between the two modes. Always profile/benchmark against a release (AOT) build when making a real performance decision for a shipped app; a debug-mode measurement is for iteration speed, not for performance conclusions.

8. Allocations in hot loops

A hot loop is any loop that runs a huge number of times relative to the rest of your program — the "hot" refers to how much CPU time is spent there, not temperature. Every new object created inside a hot loop is extra work for both the allocator (finding space) and, later, the garbage collector (reclaiming that space once it's unreachable — see D13). Reusing one mutable object instead of allocating a fresh one every iteration can remove that overhead entirely for the reused object.

Reuse vs allocate, and the ring buffer

1 object vs one per offset; the ring keeps ONE fixed array and overwrites the oldest value when full.

Input size → what is feasible: n ≤ 107 offsets → both loops are fine in time; the reuse version allocates 1 object, the immutable one allocates n (107 short-lived objects for the garbage collector).

String building is a very common accidental hot loop: repeatedly using += on a String inside a loop, because Dart strings are immutable — every += allocates an entirely new string and copies everything accumulated SO FAR into it, so total copying work across n concatenations is 1+2+...+n, which is O(n²). A StringBuffer (introduced in D04) collects fragments internally and builds the final string exactly once, doing O(n) total work — this equivalence (same output, different total work) is tested directly in verify/d20.dart, never asserted via timing.

Counting the copying: += vs StringBuffer

n(n+1)/2 characters copied by += versus n writes plus one final copy.

Input size → what is feasible: n ≤ 103 pieces → either way; n = 2·105 pieces of 10 characters → += copies ≈ 2·1011 characters (too slow), StringBuffer copies ≈ 2·106.

"Reduce allocations" does not mean "never allocate anything" — it means don't allocate NEEDLESSLY, repeatedly, inside a loop that runs many times, when a single reusable object (or a fixed-size buffer like the ring buffer above) would do the same job. Allocating a handful of objects once, outside a loop, is completely normal and not a performance concern.

Two more allocation-related tools worth knowing: typed data (Uint8List, Float64List, etc., covered in D16) stores numbers packed contiguously in a fixed-size buffer instead of as individually-boxed objects in a general List — useful for large numeric buffers (audio/image/binary data) where per-element object overhead would otherwise add up. And @pragma('vm:prefer-inline') is a VM-specific hint you CAN place above a small, extremely hot function to suggest inlining it at its call sites, removing call overhead — but this is a narrow, implementation-specific hint that the compiler is free to ignore, it can make code SLOWER if misapplied (larger inlined code hurts instruction-cache behaviour), and it should only ever follow real profiling evidence that a specific tiny function's call overhead actually matters, never be added speculatively.

Typed data: same answer, smaller storage

The inline hint compiles and changes nothing about the answer

Σ i² = n(n+1)(2n+1)/6 checked for n = 999.

Measured (M3 Pro, Dart 3.11.5, 105 elements, median of 9): summing a List<int> took about 54–56 µs and an Int32List about 40–43 µs, a ratio of about 1.3–1.4×. This was with an indexed loop (for (var i = 0; i < n; i++) s += xs[i]); on the same machine a for (final v in typedList) loop was slower, not faster, than the same loop over a List<int>, so loop style can matter as much as the type. The big win of typed data is memory and I/O layout, not loop speed.

9. const, collections, dynamic, and isolates

A const constructor call (e.g. const SizedBox(height: 8)) is canonicalized at compile time (see D13's identity lesson) — every textually-identical const expression reuses the exact same object, so a Flutter widget tree that reuses const widgets can skip re-creating (and, for widgets, potentially re-rendering) them. The exact rebuild-avoidance behaviour is a Flutter framework detail, hedged here deliberately — the language-level guarantee is only about object identity/canonicalization, not about frame timing.

const gives one shared object

Choosing the right collection (D08) for your access pattern matters more than almost any micro-optimisation: a List gives O(1) index access but O(n) "is this value present" checks; a Set/Map gives average O(1) membership/lookup by hashing, at the cost of not preserving a meaningful numeric index; a LinkedHashMap/LinkedHashSet keeps insertion order alongside hash-based lookup. Using .contains() on a List inside a loop (turning an intended O(n) algorithm into O(n²)) is one of the single most common real-world performance bugs — see the Hard-tier dedupe question below for exactly this fix.

Counting == calls: List.contains vs Set.contains

Worst-case lookup of the last element: n comparisons for a list, 1 for a hash set.

Measured (5000 lookups in a 5000-element collection, median): List.contains about 6900–7200 µs versus Set.contains about 13–28 µs, a ratio of about 250–540×. As a queue, removeAt(0) on a 2·104-element list took about 110 ms versus about 0.22–0.26 ms with ListQueue.removeFirst, about 420–530×: never use removeAt(0) as a queue.

A chain like xs.where(...).map(...).toList() uses lazy iterables — each stage doesn't build an intermediate list, it produces values on demand as the next stage asks for them. This avoids allocating throwaway intermediate lists, which is good, but chaining many such stages does add some per-element call/wrapper overhead compared to one hand-written loop doing the same filtering and mapping directly; for a hot loop processing millions of elements this is worth knowing, but for everyday code the clarity of .where().map() is usually worth far more than any difference here, which is typically small and always workload/version dependent.

Lazy chains: counting what actually runs

Measured (105 elements, filter even then double then sum): a hand-written loop took about 70–95 µs, the where().map().fold() chain about 600–715 µs, so the chain was about 7–9× slower on this workload: clarity usually wins, but not inside a hot loop over millions of elements.

Using dynamic turns off Dart's compile-time type checking for that value, and a call through a dynamic-typed reference generally cannot use the same fast, statically-resolved dispatch a concretely-typed call can — it's typically somewhat slower, though exactly how much depends on the VM/compiler version and is not something to hard-code a number for. The much bigger cost of overusing dynamic is usually correctness, not speed: you lose the compiler's help catching a typo'd method name or a wrong-type argument until runtime.

dynamic: the typo is found at run time

Measured (indexing a 105-element List<int>): through a dynamic variable 1.0× (no measurable difference, about 61–70 µs both ways), because the receiver is always the same type and the VM specialises the call. The cost of dynamic is the lost compile-time check, not a guaranteed slowdown.

For genuinely CPU-heavy work (not I/O — I/O already doesn't block Dart's event loop, see D14), moving it into a separate isolate (D15) is the real tool: it lets the main isolate's event loop stay responsive to input/UI/other work while the heavy computation runs elsewhere. This doesn't make the underlying algorithm any faster by itself (same Big-O, same total work) — it changes WHERE that work happens so it doesn't block anything else, and can add real parallelism if you split the work across several isolates on a multi-core machine.

Isolate.run: same work, same answer, different place

10. Flutter-specific notes (brief, hedged)

Flutter's own performance story deserves a full lesson of its own; briefly: a rebuild (Flutter re-running a widget's build() method) is generally cheap by itself, but doing expensive work directly inside build() multiplies that cost by every rebuild. A const widget constructor lets Flutter skip rebuilding that specific subtree when its parent rebuilds, since the framework can tell nothing about it changed. Jank means a visibly dropped/late frame — commonly described as a frame that misses its budget (roughly 16ms for a smooth 60Hz display, less on higher refresh-rate screens), causing visible stutter; the exact budget depends on the display's actual refresh rate and is best treated as "keep each frame's work small," not a hard-coded constant to test against in this lesson's non-UI verify file.

The frame budget is just arithmetic

11. Caching, memoization & precomputation

Memoization is a sticky note taped to the front of an expensive pure function: "already computed f(4)? It was 16 — here, don't redo the work." It only works safely for a pure function — one that always returns the same output for the same input and has no other effects — because the sticky note is trusted blindly on every repeat call.

The classic teaching example is a naive recursive Fibonacci, which recomputes the same smaller sub-problems an exploding number of times; memoizing it (caching each n's result the first time it's computed) turns an exponential blow-up into roughly one computation per distinct n:

Naive vs memoized: exact call counts and the closed form

calls(n) = 1 + calls(n−1) + calls(n−2) solves to 2·fib(n+1) − 1; memoized needs 2n−1 calls for n ≥ 2.

Input size → what is feasible: n ≤ 35 → naive fib is ≈ 3·107 calls (fits); n = 50 → ≈ 4·1010 calls (too slow); the memoized version handles n ≤ 90 (the result fits 64 bits) with recursion depth ≤ 91. Measured (fib(25), median): naive about 310–556 µs, memoized about 0.9–2.2 µs, about 240–390×.

Precomputation is memoization's cousin: instead of caching answers lazily as they're first asked for, you compute a useful structure UP FRONT, once, so every later query is cheap. The prefix-sum technique in this lesson's Medium-tier "optimize this loop" question below is a clean example: computing a running-total array once in O(n) turns every subsequent range-sum query into O(1), instead of re-summing the range from scratch every time.

Prefix sums: the formula as code, with the 1-indexed → 0-indexed shift

Math form: sum(l..r) = P[r] − P[l−1] (1-indexed). Dart form: prefix[hi+1] − prefix[lo] with lo = l−1, hi = r−1.

Input size → what is feasible: n, q ≤ 2·105 → re-summing is n·q = 4·1010 steps (too slow); prefix sums are n + q = 4·105.

A general-purpose cache extends memoization with real-world bounds: it can't grow forever (an LRU — Least Recently Used — policy evicts the entry that hasn't been touched in the longest time once the cache is full) and entries can go stale (a TTL — Time To Live — expires an entry after a set duration even if it's never evicted for space). The Hard-tier "cache with TTL + LRU" question below implements and tests both policies together, using an injectable fake clock so the TTL behaviour is fully deterministic and never depends on a real sleep() or wall-clock timing.

TTL + LRU with a fake clock

Never memoize a function with side effects, or one whose result can legitimately change for the same input over time (e.g. "current time", "random number", "read this mutable global") — the cache will confidently return a stale or simply wrong answer, and the bug will look like it's somewhere else entirely.

What goes wrong when you memoize an impure function

Sometimes the biggest win isn't a faster loop at all — it's an algorithmic one: picking a fundamentally better algorithm or data structure for the problem (see the Algorithms track for the deep version of this: sorting, searching, graph algorithms, dynamic programming). No amount of micro-optimisation or caching rescues an algorithm with the wrong Big-O for its input size.

Quiz

Interview questions

Cheat sheet

TermMeaning
build_runnerOrchestrates the build: scans source files, dispatches matching ones to registered builders (e.g. json_serializable, freezed)
part 'x.g.dart';Promises a generated file will be merged into this library; names must match exactly
macros (Dart)Experimental compiler-integrated codegen feature — officially discontinued January 2025; not used in current Dart
Big-ODescribes how work GROWS with input size; dominates over constant-factor micro-optimisations for large enough input
JITCompiles to machine code while running, with a warm-up phase; used in debug, dart run and hot reload
AOTCompiles to machine code before running, no warm-up; used in Flutter release builds
Hot loopA loop that runs enough times that per-iteration cost (e.g. allocations) matters
StringBufferBuilds a string in O(n) total work; avoids the O(n²) cost of repeated +=
MemoizationCaching a pure function's results by input, lazily, on first call
PrecomputationBuilding a useful structure up front so later queries are cheap (e.g. prefix sums)
LRU / TTL cacheEvicts the least-recently-used entry when full / expires entries after a fixed duration
Performance budgetA deterministic, operation-count-based regression test — never a wall-clock assertion