Testing, Debugging & DevTools Profiling
By the end of this lesson you will be able to write real package:test tests (and know exactly what each matcher checks), spot the two most common test bugs before they bite you in code review, make time- and network-dependent code testable with dependency injection, design a fake HTTP client and a fake clock, write a property-based check that hunts for bugs a hand-picked example would miss, read a stack trace and step through a buggy loop with a debugger, and read a CPU flame chart without guessing.
1. Why test? The test pyramid
Automated tests exist to answer one question fast and repeatedly: "did my last change break anything that used to work?" Without them, every change is a gamble that only gets resolved when a human (ideally you, worst case a user) happens to notice the breakage. The earlier a bug is caught, the cheaper it is to fix — a typo caught by a test takes seconds to fix; the same typo found by a customer can take hours of investigation plus the cost of their bad experience.
In Dart — a regression test passes today and fails the moment a change breaks the behaviour:
Not all tests are equal in cost and confidence. The test pyramid is a rule of thumb for how many of each kind to write:
- Unit tests (the wide base — write LOTS) — test one function or class in isolation, no real database/network/UI. Milliseconds each, pinpoint exactly what broke.
- Widget tests (the middle — write a moderate number) — in Flutter, render one widget/screen in a simulated environment and interact with it (tap, scroll, check what's on screen), without a real device.
- Integration tests (the narrow top — write a FEW) — run the whole app (or a large slice of it) end-to-end, often on a real or simulated device. Slow and more brittle, but catch problems unit tests structurally cannot see (wiring between real pieces, real platform behavior).
In Dart — the pyramid as measured numbers (cost per kind of test):
package:test, test doubles, and most debugging/profiling skills live.2. package:test: test, group, expect, matchers, setUp/tearDown, async, TDD
test() as writing one line in an inspection checklist ("check the door closes"), group() as a labeled section of the checklist ("Doors"), and expect(actual, matcher) as the actual pass/fail check ("does it match 'closes with a click'? yes/no"). setUp/tearDown are "reset the room to a known state before every check, and tidy up after" — so checklist item #7 never accidentally depends on what item #3 left behind.The real, standard tool for this in Dart is package:test, added with dart pub add dev:test and imported as import 'package:test/test.dart';. Its shape is:
import 'package:test/test.dart';
void main() {
group('isAnagram', () {
test('same letters, different order -> true', () {
expect(isAnagram('listen', 'silent'), equals(true));
});
test('different lengths -> false', () {
expect(isAnagram('abc', 'ab'), equals(false));
});
});
}
In Dart — the same shape, run through the offline harness (reporter output shown):
verify/d18.dart cannot literally dart pub add dev:test. Everywhere this page shows package:test-style code, verify/d18.dart actually runs it through a small hand-written test()/group()/setUp()/tearDown()/expect() harness (pure Dart SDK, no imports beyond dart:async/dart:math) that mirrors the real API closely enough that every example genuinely executes and is genuinely checked. In your own projects, use the real package:test — this harness is a stand-in for one offline environment, not a replacement.Common matchers you pass as the second argument to expect:
| Matcher | Passes when… | Example |
|---|---|---|
equals(x) | deep-equal to x (element-by-element for Lists) | expect([1,2,3], equals([1,2,3])); |
isA<T>() | the actual value's runtime type is T | expect(FormatException('x'), isA<FormatException>()); |
throwsA(matcher) | calling the actual (a zero-arg function) throws something matching matcher | expect(() => int.parse('abc'), throwsA(isA<FormatException>())); |
closeTo(x, delta) | a number is within delta of x — essential for floating point | expect(0.1 + 0.2, closeTo(0.3, 1e-9)); |
contains(x) | a String contains substring x, or an Iterable contains element x | expect('hello world', contains('wor')); |
All five are verified working (against our harness) in verify/d18.dart's registerMatcherDemoTests().
In Dart — all five matchers, plus the message a failing matcher prints:
expect(0.1 + 0.2, equals(0.3)) FAILS in Dart (and every language using IEEE-754 doubles) because 0.1 + 0.2 is actually 0.30000000000000004, not exactly 0.3. This is exactly why closeTo exists — never compare floating-point numbers with plain equality.In Dart — why closeTo exists:
The frames above trace the classic TDD loop, red → green → refactor: (1) red — write a test for behavior that doesn't exist yet (or is a stub), watch it fail; (2) green — write the smallest code that makes it pass; (3) refactor — clean the code up with the test suite as a safety net, confident that if you break something the tests will tell you immediately.
In Dart — red, green, and the tests that stay unchanged:
setUp runs before every test in its group (or file); tearDown runs after every test, whether it passed or failed. They exist so tests don't leak state into each other:
In Dart — the exact order of setUp / tests / tearDown:
setUp (e.g. as a top-level variable initialized once) instead of inside it. Then test order starts to matter — a test that mutates the shared object can silently make a LATER test fail (or worse, silently pass when it shouldn't) depending on what ran before it. Always build fresh state in setUp.In Dart — the shared-state pitfall made visible:
Async code is tested by making the test body itself async and await-ing whatever you're testing, exactly like normal async code:
test('fetchValue resolves to 42', () async {
final v = await fetchValue();
expect(v, equals(42));
});
In Dart — an async test that awaits:
In Dart — PASS is printed before the real check runs:
async/await on an async test is one of the single most common real-world test bugs, and it is dangerous precisely because the test still reports green. If the test body isn't async and never returns or awaits the Future it kicked off, the test runner considers the test finished the instant the (synchronous) body returns — long before the real work, and the expect inside it, ever run. verify/d18.dart proves this live: registerMissingAwaitBuggyTest() is reported PASSED, and only after an explicit delay does the real (deliberately wrong) mismatch actually occur — too late for anyone to see it fail.async but that never awaits the thing it calls — Dart won't complain (an unawaited Future is legal), and the analyzer's unawaited_futures lint is the main defense; that's why the buggy example above is annotated // ignore: unawaited_futures in verify/d18.dart — normally you would NOT silence that warning, you'd fix the bug it's pointing at.expect that never actually ran (because of a missing await, or because it's inside a branch that was never reached) contributes nothing, and the test suite cannot tell the difference from your build output alone.Input size → what's feasible: a unit test of a pure function takes ≈ 1 ms, so 104 of them run in about 10 s and a 105-test suite needs parallel runs; one 3 s end-to-end test is as costly as 1500 unit tests.
3. Test doubles: fakes, mocks, stubs & dependency injection
Three words get used loosely, but they mean different things:
- Stub — a bare-minimum replacement that returns canned answers and does nothing else. "When asked for the time, always say 9:00am."
- Fake — a WORKING, simplified implementation of the real thing, built by hand.
FakeHttpClientbelow is a fake: it really stores and returns data, just from aMapinstead of a network. - Mock — an object created by a mocking FRAMEWORK (in Dart:
mockitoor the more modern, null-safety-firstmocktail) that additionally lets you assert HOW it was called: "wassave()called exactly once, with these arguments?" Mocks verify interactions; fakes and stubs just provide data.
mocktail (the currently favored choice for new Dart/Flutter code, because it needs no code generation) lets you write class MockHttpClient extends Mock implements HttpClient {}, then when(() => mock.get(any())).thenAnswer((_) async => 'Asha'); to stub a return value, and later verify(() => mock.get('/users/7')).called(1); to assert it was actually invoked. mockito is the older, still-common alternative, historically requiring build_runner code generation for null-safe mocks: annotate a source file with @GenerateMocks([HttpClient]) (or @GenerateNiceMocks([MockSpec<HttpClient>()]) for mocks that don't throw on unstubbed calls), run dart run build_runner build, and it emits a companion *.mocks.dart file with the generated mock classes — though mockito also has a manual mode (hand-writing a class that extends Mock) for simple cases. Neither package is installed in this offline environment (pub.dev is unreachable here), so this page describes their real API accurately but verify/d18.dart demonstrates the underlying IDEA with a hand-written fake instead, which needs no package at all.In Dart — stub, fake and mock side by side:
None of this works unless the code being tested is willing to ACCEPT a replacement — which means it must not construct its real dependency internally. This is dependency injection (DI): pass the dependency IN (as a constructor or function parameter) instead of reaching out and creating/looking it up yourself.
In Dart — the injected clock makes every answer repeatable:
DateTime.now(), Random(), or a concrete http.Client() directly, buried inside a method, CANNOT be tested deterministically — every test run depends on the real wall clock, real randomness, or a real network. The fix is always the same shape: define a small interface (Clock, HttpClient), have the REAL implementation wrap the real thing, inject the interface, and write a FAKE for tests.In Dart — using the fake HTTP client (no network):
Input size → what's feasible: a fake or clock is O(1) per call, so 106 injected calls ≈ 2·107 steps is instant; a real network call is ≈ 108 steps of waiting each, which is why unit tests must not make them.
4. Writing good tests: AAA, edge cases, property-based testing
The standard shape is Arrange – Act – Assert (AAA): build the inputs and any fakes (Arrange), call the one thing under test (Act), then check the outcome (Assert) — usually as three visually separate chunks, even without comments:
test('isExpired: token has expired', () {
final issued = DateTime(2026, 1, 1); // Arrange
final clock = FakeClock(DateTime(2026, 1, 2));
final expired = isExpired(issued, const Duration(hours: 12), clock); // Act
expect(expired, equals(true)); // Assert
});
In Dart — the AAA test above, executed:
Beyond the "happy path", deliberately test the edges where bugs hide: empty input, exactly one element, a large/typical input, and the boundary values right at a condition's cutoff:
| Edge case | Why it matters | isPalindrome example |
|---|---|---|
| Empty | loops/indexing that assume ≥1 element often break silently or throw | isPalindrome('') → true (vacuously — verified) |
| One element | often the smallest input that still enters a loop body | isPalindrome('a') → true (verified) |
| Many / typical | the "normal" case — necessary but never sufficient on its own | isPalindrome('level') → true (verified) |
| Boundary | off-by-one bugs live exactly at < vs <=, n vs n-1 | even- vs odd-length strings meet in the middle differently — both verified |
| Invalid / hostile | real input includes punctuation, case, whitespace an author didn't picture | isPalindrome('A man, a plan, a canal: Panama') → true (verified) |
In Dart — every edge-case row, executed:
Hand-picked examples only catch the bugs the AUTHOR thought of. Property-based testing flips this: instead of writing "does it sort [3,1,2] correctly?", you write a general PROPERTY that should hold for ANY input — "the output equals what a trusted reference implementation produces" — and let the computer generate many random inputs and hunt for a counterexample.
In Dart — the property check on the page's exact seed:
package:glados in the Dart ecosystem, or QuickCheck in Haskell where the idea originated) also do shrinking: when a random input fails, they automatically try smaller/simpler variants of it until they find the smallest input that still fails — which is the exact same idea as the delta-debugging technique in section 6. verify/d18.dart demonstrates the core PROPERTY-CHECKING idea by hand with a seeded Random (for reproducibility) rather than pulling in that package.Input size → what's feasible: a property test with 104 random arrays of length ≤ 100 costs about 107 steps (well under 1 s); length 105 with a quadratic candidate would be 1010 steps, so keep random inputs small and many.
5. Coverage, and its limits
With the real package:test installed, you generate coverage with:
dart test --coverage=coverage dart pub global activate coverage # one-time setup dart pub global run coverage:format_coverage \ --lcov --in=coverage --out=coverage/lcov.info \ --packages=.dart_tool/package_config.json --report-on=lib
which produces an lcov.info file that editors (VS Code's "Coverage Gutters", for example) and CI dashboards render as a percentage and a line-by-line highlight of what ran and what didn't.
In Dart — 100% coverage with a bug still alive:
await bug — the buggy test still "covers" the code inside the .then callback, it just never verifies the value in time to matter). Coverage also can't see MISSING code paths that were never written at all (a case you forgot to handle entirely leaves no line to "not cover"). Treat coverage as a tool for finding UN-tested code to look at, never as a certificate of correctness.dart test --coverage measures which lines RAN, not which behaviors were actually CHECKED. Use it to find gaps, not as a finish line — a thoughtful edge-case test beats a high coverage number with weak assertions every time.Input size → what's feasible: coverage instrumentation slows a run by roughly 2× to 10×, fine for a 104-test suite; it measures lines run, never whether the right value was checked.
6. Debugging: stack traces, breakpoints, assert, debugger()
print statements is like investigating a break-in by asking a few witnesses "did you see anything?" one at a time and writing down whatever they happen to volunteer. A real debugger is like pausing time at the scene and being able to walk around, open drawers, and inspect anything you want — far more thorough, and it doesn't require guessing in advance which detail will turn out to matter.Reading a stack trace (covered in depth in D12) is debugging step zero: #0 is the exact throw site, each higher number is one caller further out — read top to bottom and start at #0.
In Dart — reading frames, #0 first:
print vs a real debugger: print is fast for a single known question ("what is x right here?") but requires editing code, re-running, and guessing in advance what to print. A debugger with breakpoints lets you pause execution at any line and inspect EVERY variable in scope, without changing the code, then choose one of three ways to keep going:
- Step over — run the current line (including any function it calls) and stop at the next line in the SAME function. Use this when you trust the function being called.
- Step into — jump INSIDE the function call on the current line, to see what it does internally. Use this when you suspect the bug is inside that call.
- Step out — run until the current function returns, then stop in its caller. Use this once you've learned what you needed from inside a function and want to get back out.
In Dart — step over versus step into, as the lines each one visits:
In Dart — what the watch panel shows on the off-by-one loop:
assert(condition, 'message') (from D12) is a lightweight, debug-only internal consistency check — it throws AssertionError when condition is false, but ONLY when assertions are enabled (debug mode / --enable-asserts); it's stripped entirely from release builds and plain dart run. import 'dart:developer'; debugger(); is a different, stronger tool: it's a programmatic breakpoint — when the isolate is already being watched by an attached debugger (DevTools, an IDE), execution pauses right there, exactly as if you'd clicked to set a breakpoint on that line. It is meant to be reached from CODE you're actively debugging with tooling attached, not left in code paths that might run unattended (outside of a real debugging session its behavior depends on the tooling around it, so treat it as a temporary debugging aid you remove afterward, not a permanent code path).
In Dart — assert versus real validation, and debugger():
print, package:logging gives you structured, leveled logs: final log = Logger('MyClass'); log.info('started'); log.warning('retrying'); log.severe('failed', error, stackTrace);, with a single place to configure what's shown (Logger.root.level = Level.ALL; Logger.root.onRecord.listen((r) => print('${r.level.name}: ${r.message}'));). Unlike scattered print calls, logs carry a level (so you can silence noisy ones in production) and a named source (so you know WHICH part of the app logged it) — this page shows the real API; it is not installed in this offline environment, so it is illustrative, not executed here.In Dart — leveled logging with one threshold (plain-Dart stand-in for package:logging):
A very different debugging technique applies once you know code USED to work and now doesn't, but not which commit broke it: binary search over history, exactly what git bisect automates. Instead of checking every commit one by one, jump to the MIDDLE of the suspect range, test it (ideally with an automated test), and keep only the half that still contains the bug — halving the search space every time.
In Dart — how many checks git bisect needs:
git bisect start, then git bisect bad (mark the current, broken commit) and git bisect good <commit> (mark a known-working one) puts Git into exactly this binary-search mode: it checks out the middle commit for you, you run your test (or git bisect run <script> to automate it entirely) and answer good/bad, and Git narrows the range until only one commit remains — the exact one that introduced the bug.In Dart — delta debugging (shrinking a failing input), run on the lesson's numbers:
#0. Use step over/into/out deliberately based on whether you trust the code being stepped through. assert is debug-only; debugger() is a code-level breakpoint for an ALREADY-attached debugger. Binary search (manually, or via git bisect) turns "which of 500 commits broke this" into about 9 steps (log₂ 500 ≈ 9).Input size → what's feasible: bisecting n = 500 commits needs ⌈log2 500⌉ = 9 checks, n = 106 commits only 20, while checking every commit would take 500 or 106 test runs.
7. Profiling with DevTools: flame charts, memory, Stopwatch pitfalls
Dart/Flutter's DevTools (launched from your IDE, or via dart devtools) gives you three profiling views relevant here:
- CPU profiler — records which functions were executing, how often, and for how long, then draws it as a flame chart: nested horizontal bars, one row per call-stack depth, bar WIDTH = time spent.
- Memory view — tracks how much memory is allocated over time and lets you take heap snapshots to see what kinds of objects are piling up (the classic symptom of a memory leak: a steadily climbing line that never drops back down after garbage collection).
- Timeline / performance view — shows frame-by-frame UI rendering cost in a Flutter app, flagging "janky" frames that took too long to build/render to hit 60/120fps.
A flame chart reads like a call-stack snapshot repeated over time, left to right = time, and each bar's CHILDREN (the row below it) are the functions IT called:
In Dart — self time versus total time:
StreamSubscription or AnimationController never .cancel()/.dispose()d, or a long-lived object (a singleton, a cache) holding a closure that captures a short-lived widget's BuildContext. The memory view's heap snapshot, taken before and after an action that SHOULD free memory, is how you confirm a leak: if the "before" and "after" object counts for a type don't drop as expected, something is still holding a reference.In Dart — a subscription nobody cancels:
For quick, code-level timing without opening DevTools, Dart's Stopwatch gives microsecond-resolution timing — but a naive benchmark is easy to get wrong:
In Dart — a measured benchmark with warm-up and a sink:
verify/d18.dart's benchmark() helper: (1) JIT warm-up — the first runs of hot code are slower because the VM hasn't yet compiled it to optimized native code; time only AFTER a warm-up period. (2) Dead-code elimination — if a benchmarked function's result is never used for anything observable, an optimizing compiler is technically permitted to notice the whole loop has no effect and remove it, so you measure "how fast is nothing" — always feed results into something that's used later (a "sink" variable that gets printed/returned/checked), as benchmark() does with _sink.Stopwatch micro-benchmarks only for narrow "is approach A or B faster for this exact operation" questions — and even then, run each approach many times and compare distributions, not single readings, because system noise (other processes, thermal throttling, GC pauses) makes any single measurement unreliable.Stopwatch benchmark always warms up first and always uses its result via a sink, or the JIT and the optimizer will quietly lie to you.Input size → what's feasible: a micro-benchmark needs ≈ 103 to 106 timed calls to beat the 1 µs clock resolution; a body of ~100 steps at 106 calls = 108 steps ≈ 1 s, so size iterations to that budget.
Quiz
Interview questions
Cheat sheet
| Concept | Syntax / rule |
|---|---|
| Test pyramid | Many unit tests (fast, precise) → some widget tests → few integration tests (slow, end-to-end) |
test / group | test('desc', () { ... }); registers one test; group('name', () { ...tests... }) labels/nests several |
| Matchers | equals(x), isA<T>(), throwsA(m), closeTo(x, delta), contains(x) |
setUp / tearDown | run before / after EVERY test in scope — build fresh state here, never share a mutable object across tests |
| Async test | make the test body async and await what you're testing — a missing await reports green with no real check having happened yet |
| TDD loop | red (failing test) → green (minimal fix) → refactor (clean up, safety net already in place) |
| Stub / Fake / Mock | Stub = canned answer. Fake = real, simplified working implementation. Mock = framework double (mockito/mocktail) that also verifies HOW it was called. |
| Dependency injection | pass a Clock/HttpClient/etc. IN, instead of constructing it inside the function — the only way to make time/network/randomness testable |
| AAA | Arrange (build inputs/fakes) → Act (call the one thing under test) → Assert (check the outcome) |
| Edge cases | always test: empty, one, many, boundary, invalid/hostile input |
| Property-based test | check a general property ("matches a trusted reference") against MANY random inputs, not a few hand-picked ones |
| Coverage | dart test --coverage=coverage — measures which LINES ran, not which behaviors were actually CHECKED; 100% ≠ bug-free |
| Stack trace | #0 = throw site (innermost); read top to bottom, each higher number is one caller further out |
| Step over/into/out | over = stay in this function, skip past a call; into = go inside the call; out = finish this function, land in its caller |
assert vs debugger() | assert = debug-only internal check, stripped in release. debugger() (dart:developer) = programmatic breakpoint for an already-attached debugger. |
| git bisect | binary search over commit history: git bisect start, bisect bad, bisect good <c>, or automate with bisect run <script> |
| Flame chart | bar width = time; nested bars = callees; look at SELF time (own code only), not just total width, to find the real hotspot |
Stopwatch benchmarking | warm up before timing (JIT), and feed every result into a used "sink" (avoid dead-code elimination) |