Testing, Debugging & DevTools Profiling

By the end of this lesson you will be able to write real package:test tests (and know exactly what each matcher checks), spot the two most common test bugs before they bite you in code review, make time- and network-dependent code testable with dependency injection, design a fake HTTP client and a fake clock, write a property-based check that hunts for bugs a hand-picked example would miss, read a stack trace and step through a buggy loop with a debugger, and read a CPU flame chart without guessing.

1. Why test? The test pyramid

A test is a tiny robot inspector you build once and can run forever. A human tester checks your app by hand and gets bored, tired, and inconsistent after the tenth click-through. A test never gets bored: it checks the exact same thing, the exact same way, in milliseconds, every single time you change a line of code — and it tells you the SECOND you break something, instead of a user finding out three weeks later in production.

Automated tests exist to answer one question fast and repeatedly: "did my last change break anything that used to work?" Without them, every change is a gamble that only gets resolved when a human (ideally you, worst case a user) happens to notice the breakage. The earlier a bug is caught, the cheaper it is to fix — a typo caught by a test takes seconds to fix; the same typo found by a customer can take hours of investigation plus the cost of their bad experience.

In Dart — a regression test passes today and fails the moment a change breaks the behaviour:

Not all tests are equal in cost and confidence. The test pyramid is a rule of thumb for how many of each kind to write:

In Dart — the pyramid as measured numbers (cost per kind of test):

The pyramid shape is about a trade-off, not a law: unit tests are fast and precise but can all pass while the pieces still don't fit together when wired for real; integration tests catch that wiring but are slow, flakier, and a failure tells you much less about WHERE the bug is. Most teams aim for roughly 70% unit, 20% widget/component, 10% integration — few enough slow tests that the whole suite still runs often, enough fast ones that most bugs are caught precisely.
An "ice cream cone" anti-pattern — mostly slow, flaky end-to-end tests and almost no unit tests — is a common mistake in real codebases. It feels productive (you're testing "the real thing") but the suite becomes slow to run, painful to debug (a failure could be anywhere), and flaky, so people start ignoring red builds — which defeats the entire purpose of having tests.
Tests exist to catch regressions early and cheaply. Favor many fast, precise unit tests; some widget tests; a few slow, high-confidence integration tests. This whole lesson focuses mostly on the unit-test layer, because that's where package:test, test doubles, and most debugging/profiling skills live.

2. package:test: test, group, expect, matchers, setUp/tearDown, async, TDD

Think of test() as writing one line in an inspection checklist ("check the door closes"), group() as a labeled section of the checklist ("Doors"), and expect(actual, matcher) as the actual pass/fail check ("does it match 'closes with a click'? yes/no"). setUp/tearDown are "reset the room to a known state before every check, and tidy up after" — so checklist item #7 never accidentally depends on what item #3 left behind.

The real, standard tool for this in Dart is package:test, added with dart pub add dev:test and imported as import 'package:test/test.dart';. Its shape is:

import 'package:test/test.dart';

void main() {
  group('isAnagram', () {
    test('same letters, different order -> true', () {
      expect(isAnagram('listen', 'silent'), equals(true));
    });
    test('different lengths -> false', () {
      expect(isAnagram('abc', 'ab'), equals(false));
    });
  });
}

In Dart — the same shape, run through the offline harness (reporter output shown):

This environment cannot reach pub.dev, so verify/d18.dart cannot literally dart pub add dev:test. Everywhere this page shows package:test-style code, verify/d18.dart actually runs it through a small hand-written test()/group()/setUp()/tearDown()/expect() harness (pure Dart SDK, no imports beyond dart:async/dart:math) that mirrors the real API closely enough that every example genuinely executes and is genuinely checked. In your own projects, use the real package:test — this harness is a stand-in for one offline environment, not a replacement.

Common matchers you pass as the second argument to expect:

MatcherPasses when…Example
equals(x)deep-equal to x (element-by-element for Lists)expect([1,2,3], equals([1,2,3]));
isA<T>()the actual value's runtime type is Texpect(FormatException('x'), isA<FormatException>());
throwsA(matcher)calling the actual (a zero-arg function) throws something matching matcherexpect(() => int.parse('abc'), throwsA(isA<FormatException>()));
closeTo(x, delta)a number is within delta of x — essential for floating pointexpect(0.1 + 0.2, closeTo(0.3, 1e-9));
contains(x)a String contains substring x, or an Iterable contains element xexpect('hello world', contains('wor'));

All five are verified working (against our harness) in verify/d18.dart's registerMatcherDemoTests().

In Dart — all five matchers, plus the message a failing matcher prints:

expect(0.1 + 0.2, equals(0.3)) FAILS in Dart (and every language using IEEE-754 doubles) because 0.1 + 0.2 is actually 0.30000000000000004, not exactly 0.3. This is exactly why closeTo exists — never compare floating-point numbers with plain equality.

In Dart — why closeTo exists:

The frames above trace the classic TDD loop, red → green → refactor: (1) red — write a test for behavior that doesn't exist yet (or is a stub), watch it fail; (2) green — write the smallest code that makes it pass; (3) refactor — clean the code up with the test suite as a safety net, confident that if you break something the tests will tell you immediately.

In Dart — red, green, and the tests that stay unchanged:

TDD's real value isn't "write tests first" as a ritual — it's that watching a NEW test fail (red) proves the test can actually detect the bug it's meant to catch, before you trust it to protect you during every future refactor.

setUp runs before every test in its group (or file); tearDown runs after every test, whether it passed or failed. They exist so tests don't leak state into each other:

In Dart — the exact order of setUp / tests / tearDown:

A common bug is sharing ONE mutable object across tests by creating it outside setUp (e.g. as a top-level variable initialized once) instead of inside it. Then test order starts to matter — a test that mutates the shared object can silently make a LATER test fail (or worse, silently pass when it shouldn't) depending on what ran before it. Always build fresh state in setUp.

In Dart — the shared-state pitfall made visible:

Async code is tested by making the test body itself async and await-ing whatever you're testing, exactly like normal async code:

test('fetchValue resolves to 42', () async {
  final v = await fetchValue();
  expect(v, equals(42));
});

In Dart — an async test that awaits:

In Dart — PASS is printed before the real check runs:

Forgetting async/await on an async test is one of the single most common real-world test bugs, and it is dangerous precisely because the test still reports green. If the test body isn't async and never returns or awaits the Future it kicked off, the test runner considers the test finished the instant the (synchronous) body returns — long before the real work, and the expect inside it, ever run. verify/d18.dart proves this live: registerMissingAwaitBuggyTest() is reported PASSED, and only after an explicit delay does the real (deliberately wrong) mismatch actually occur — too late for anyone to see it fail.
A second, sneakier bug in the same family: writing a test whose body is async but that never awaits the thing it calls — Dart won't complain (an unawaited Future is legal), and the analyzer's unawaited_futures lint is the main defense; that's why the buggy example above is annotated // ignore: unawaited_futures in verify/d18.dart — normally you would NOT silence that warning, you'd fix the bug it's pointing at.
A green test tells you "nothing FAILED", not "everything was CHECKED" — an expect that never actually ran (because of a missing await, or because it's inside a branch that was never reached) contributes nothing, and the test suite cannot tell the difference from your build output alone.

Input size → what's feasible: a unit test of a pure function takes ≈ 1 ms, so 104 of them run in about 10 s and a 105-test suite needs parallel runs; one 3 s end-to-end test is as costly as 1500 unit tests.

3. Test doubles: fakes, mocks, stubs & dependency injection

Testing a smoke alarm by setting a real fire is insane — instead you use a test button that simulates the alarm condition safely and instantly. A test double is that test button for code: a stand-in for a real, slow, non-deterministic, or dangerous dependency (a network, a clock, a payment processor) that behaves predictably enough to test the code AROUND it.

Three words get used loosely, but they mean different things:

mocktail (the currently favored choice for new Dart/Flutter code, because it needs no code generation) lets you write class MockHttpClient extends Mock implements HttpClient {}, then when(() => mock.get(any())).thenAnswer((_) async => 'Asha'); to stub a return value, and later verify(() => mock.get('/users/7')).called(1); to assert it was actually invoked. mockito is the older, still-common alternative, historically requiring build_runner code generation for null-safe mocks: annotate a source file with @GenerateMocks([HttpClient]) (or @GenerateNiceMocks([MockSpec<HttpClient>()]) for mocks that don't throw on unstubbed calls), run dart run build_runner build, and it emits a companion *.mocks.dart file with the generated mock classes — though mockito also has a manual mode (hand-writing a class that extends Mock) for simple cases. Neither package is installed in this offline environment (pub.dev is unreachable here), so this page describes their real API accurately but verify/d18.dart demonstrates the underlying IDEA with a hand-written fake instead, which needs no package at all.

In Dart — stub, fake and mock side by side:

None of this works unless the code being tested is willing to ACCEPT a replacement — which means it must not construct its real dependency internally. This is dependency injection (DI): pass the dependency IN (as a constructor or function parameter) instead of reaching out and creating/looking it up yourself.

In Dart — the injected clock makes every answer repeatable:

Code that calls DateTime.now(), Random(), or a concrete http.Client() directly, buried inside a method, CANNOT be tested deterministically — every test run depends on the real wall clock, real randomness, or a real network. The fix is always the same shape: define a small interface (Clock, HttpClient), have the REAL implementation wrap the real thing, inject the interface, and write a FAKE for tests.

In Dart — using the fake HTTP client (no network):

Stub = canned answer, no logic. Fake = a real, simplified working implementation. Mock = a framework-generated double that also verifies HOW it was called. Dependency injection is what makes swapping any of them in for tests possible in the first place — inject clocks, random sources, and network/database clients instead of constructing them inside the function that uses them.

Input size → what's feasible: a fake or clock is O(1) per call, so 106 injected calls ≈ 2·107 steps is instant; a real network call is ≈ 108 steps of waiting each, which is why unit tests must not make them.

4. Writing good tests: AAA, edge cases, property-based testing

A well-structured test reads like a lab report: set up the experiment, run it once, read the result. Mixing setup, action, and checking together on every line is like a lab report that jumbles method and results into one paragraph — technically all the information is there, but nobody, including future-you, can quickly tell what's actually being tested.

The standard shape is Arrange – Act – Assert (AAA): build the inputs and any fakes (Arrange), call the one thing under test (Act), then check the outcome (Assert) — usually as three visually separate chunks, even without comments:

test('isExpired: token has expired', () {
  final issued = DateTime(2026, 1, 1);          // Arrange
  final clock = FakeClock(DateTime(2026, 1, 2));

  final expired = isExpired(issued, const Duration(hours: 12), clock); // Act

  expect(expired, equals(true));                // Assert
});

In Dart — the AAA test above, executed:

Beyond the "happy path", deliberately test the edges where bugs hide: empty input, exactly one element, a large/typical input, and the boundary values right at a condition's cutoff:

Edge caseWhy it mattersisPalindrome example
Emptyloops/indexing that assume ≥1 element often break silently or throwisPalindrome('') → true (vacuously — verified)
One elementoften the smallest input that still enters a loop bodyisPalindrome('a') → true (verified)
Many / typicalthe "normal" case — necessary but never sufficient on its ownisPalindrome('level') → true (verified)
Boundaryoff-by-one bugs live exactly at < vs <=, n vs n-1even- vs odd-length strings meet in the middle differently — both verified
Invalid / hostilereal input includes punctuation, case, whitespace an author didn't pictureisPalindrome('A man, a plan, a canal: Panama') → true (verified)

In Dart — every edge-case row, executed:

A test suite that only exercises the happy path gives false confidence — 100% of tests passing tells you nothing about the empty-list crash nobody wrote a test for. When reviewing someone else's (or your own) tests, actively ask "what INPUT would break this?" before trusting a green suite.

Hand-picked examples only catch the bugs the AUTHOR thought of. Property-based testing flips this: instead of writing "does it sort [3,1,2] correctly?", you write a general PROPERTY that should hold for ANY input — "the output equals what a trusted reference implementation produces" — and let the computer generate many random inputs and hunt for a counterexample.

In Dart — the property check on the page's exact seed:

Real property-testing libraries (e.g. package:glados in the Dart ecosystem, or QuickCheck in Haskell where the idea originated) also do shrinking: when a random input fails, they automatically try smaller/simpler variants of it until they find the smallest input that still fails — which is the exact same idea as the delta-debugging technique in section 6. verify/d18.dart demonstrates the core PROPERTY-CHECKING idea by hand with a seeded Random (for reproducibility) rather than pulling in that package.
Structure tests as Arrange–Act–Assert. Always test empty, one, many, and boundary inputs, not just the happy path. A property test ("matches a trusted reference on many random inputs") can catch bugs a handful of hand-picked examples never would.

Input size → what's feasible: a property test with 104 random arrays of length ≤ 100 costs about 107 steps (well under 1 s); length 105 with a quadratic candidate would be 1010 steps, so keep random inputs small and many.

5. Coverage, and its limits

Code coverage is a metal detector sweep of a beach, not a guarantee the beach is safe. It tells you which parts of the sand you walked over (which LINES ran during your tests) — it says nothing about whether you actually looked closely enough at what was there (whether the right ASSERTIONS ran against the right values).

With the real package:test installed, you generate coverage with:

dart test --coverage=coverage
dart pub global activate coverage         # one-time setup
dart pub global run coverage:format_coverage \
  --lcov --in=coverage --out=coverage/lcov.info \
  --packages=.dart_tool/package_config.json --report-on=lib

which produces an lcov.info file that editors (VS Code's "Coverage Gutters", for example) and CI dashboards render as a percentage and a line-by-line highlight of what ran and what didn't.

In Dart — 100% coverage with a bug still alive:

100% line coverage does NOT mean 0 bugs. A line can execute without its result ever being checked (recall section 2's missing-await bug — the buggy test still "covers" the code inside the .then callback, it just never verifies the value in time to matter). Coverage also can't see MISSING code paths that were never written at all (a case you forgot to handle entirely leaves no line to "not cover"). Treat coverage as a tool for finding UN-tested code to look at, never as a certificate of correctness.
dart test --coverage measures which lines RAN, not which behaviors were actually CHECKED. Use it to find gaps, not as a finish line — a thoughtful edge-case test beats a high coverage number with weak assertions every time.

Input size → what's feasible: coverage instrumentation slows a run by roughly 2× to 10×, fine for a 104-test suite; it measures lines run, never whether the right value was checked.

6. Debugging: stack traces, breakpoints, assert, debugger()

Debugging with only print statements is like investigating a break-in by asking a few witnesses "did you see anything?" one at a time and writing down whatever they happen to volunteer. A real debugger is like pausing time at the scene and being able to walk around, open drawers, and inspect anything you want — far more thorough, and it doesn't require guessing in advance which detail will turn out to matter.

Reading a stack trace (covered in depth in D12) is debugging step zero: #0 is the exact throw site, each higher number is one caller further out — read top to bottom and start at #0.

In Dart — reading frames, #0 first:

print vs a real debugger: print is fast for a single known question ("what is x right here?") but requires editing code, re-running, and guessing in advance what to print. A debugger with breakpoints lets you pause execution at any line and inspect EVERY variable in scope, without changing the code, then choose one of three ways to keep going:

In Dart — step over versus step into, as the lines each one visits:

In Dart — what the watch panel shows on the off-by-one loop:

The exact moment a value becomes wrong is often several lines BEFORE the crash (e.g. an index computed with the wrong loop bound, used several statements later). Stepping one line at a time from the very top of a function is slow — set the breakpoint as close as you can to where you SUSPECT the value first goes bad, then step forward from there watching the specific variable.

assert(condition, 'message') (from D12) is a lightweight, debug-only internal consistency check — it throws AssertionError when condition is false, but ONLY when assertions are enabled (debug mode / --enable-asserts); it's stripped entirely from release builds and plain dart run. import 'dart:developer'; debugger(); is a different, stronger tool: it's a programmatic breakpoint — when the isolate is already being watched by an attached debugger (DevTools, an IDE), execution pauses right there, exactly as if you'd clicked to set a breakpoint on that line. It is meant to be reached from CODE you're actively debugging with tooling attached, not left in code paths that might run unattended (outside of a real debugging session its behavior depends on the tooling around it, so treat it as a temporary debugging aid you remove afterward, not a permanent code path).

In Dart — assert versus real validation, and debugger():

For anything beyond a stray print, package:logging gives you structured, leveled logs: final log = Logger('MyClass'); log.info('started'); log.warning('retrying'); log.severe('failed', error, stackTrace);, with a single place to configure what's shown (Logger.root.level = Level.ALL; Logger.root.onRecord.listen((r) => print('${r.level.name}: ${r.message}'));). Unlike scattered print calls, logs carry a level (so you can silence noisy ones in production) and a named source (so you know WHICH part of the app logged it) — this page shows the real API; it is not installed in this offline environment, so it is illustrative, not executed here.

In Dart — leveled logging with one threshold (plain-Dart stand-in for package:logging):

A very different debugging technique applies once you know code USED to work and now doesn't, but not which commit broke it: binary search over history, exactly what git bisect automates. Instead of checking every commit one by one, jump to the MIDDLE of the suspect range, test it (ideally with an automated test), and keep only the half that still contains the bug — halving the search space every time.

In Dart — how many checks git bisect needs:

git bisect start, then git bisect bad (mark the current, broken commit) and git bisect good <commit> (mark a known-working one) puts Git into exactly this binary-search mode: it checks out the middle commit for you, you run your test (or git bisect run <script> to automate it entirely) and answer good/bad, and Git narrows the range until only one commit remains — the exact one that introduced the bug.

In Dart — delta debugging (shrinking a failing input), run on the lesson's numbers:

Read stack traces top-down from #0. Use step over/into/out deliberately based on whether you trust the code being stepped through. assert is debug-only; debugger() is a code-level breakpoint for an ALREADY-attached debugger. Binary search (manually, or via git bisect) turns "which of 500 commits broke this" into about 9 steps (log₂ 500 ≈ 9).

Input size → what's feasible: bisecting n = 500 commits needs ⌈log2 500⌉ = 9 checks, n = 106 commits only 20, while checking every commit would take 500 or 106 test runs.

7. Profiling with DevTools: flame charts, memory, Stopwatch pitfalls

If debugging is "why is this WRONG", profiling is "why is this SLOW (or eating memory)". A CPU profiler is like a factory manager standing over the assembly line with a stopwatch, timing every single station, so instead of guessing which step is the bottleneck, you get an exact, measured answer.

Dart/Flutter's DevTools (launched from your IDE, or via dart devtools) gives you three profiling views relevant here:

A flame chart reads like a call-stack snapshot repeated over time, left to right = time, and each bar's CHILDREN (the row below it) are the functions IT called:

In Dart — self time versus total time:

The widest bar is not necessarily the slowest FUNCTION — it might just be a thin wrapper that calls something slow underneath. The number you actually want is self time (time spent in that function's OWN code, excluding its children) versus total time (self time plus everything it called). A function with huge total time but tiny self time isn't the problem — look at its widest, deepest CHILD instead; a function with high SELF time, wherever it sits in the tree, is where the CPU is actually spending its cycles.
A memory leak in Dart/Flutter is almost always an object that's technically still REACHABLE (so the garbage collector correctly refuses to free it) when the programmer intended it to be gone — classic causes: a StreamSubscription or AnimationController never .cancel()/.dispose()d, or a long-lived object (a singleton, a cache) holding a closure that captures a short-lived widget's BuildContext. The memory view's heap snapshot, taken before and after an action that SHOULD free memory, is how you confirm a leak: if the "before" and "after" object counts for a type don't drop as expected, something is still holding a reference.

In Dart — a subscription nobody cancels:

For quick, code-level timing without opening DevTools, Dart's Stopwatch gives microsecond-resolution timing — but a naive benchmark is easy to get wrong:

In Dart — a measured benchmark with warm-up and a sink:

Two classic benchmarking traps, both demonstrated (and guarded against) in verify/d18.dart's benchmark() helper: (1) JIT warm-up — the first runs of hot code are slower because the VM hasn't yet compiled it to optimized native code; time only AFTER a warm-up period. (2) Dead-code elimination — if a benchmarked function's result is never used for anything observable, an optimizing compiler is technically permitted to notice the whole loop has no effect and remove it, so you measure "how fast is nothing" — always feed results into something that's used later (a "sink" variable that gets printed/returned/checked), as benchmark() does with _sink.
Micro-benchmarks measure ONE thing in isolation, under conditions (warm caches, no other work happening, a specific input size) that may not match production at all. Prefer profiling REAL workloads (the CPU profiler on an actual user flow) for "is my app slow" questions, and reach for Stopwatch micro-benchmarks only for narrow "is approach A or B faster for this exact operation" questions — and even then, run each approach many times and compare distributions, not single readings, because system noise (other processes, thermal throttling, GC pauses) makes any single measurement unreliable.
Read a flame chart's SELF time, not just bar width, to find the real hotspot. A memory leak shows up as a heap snapshot count that doesn't drop after an action that should free it. A trustworthy Stopwatch benchmark always warms up first and always uses its result via a sink, or the JIT and the optimizer will quietly lie to you.

Input size → what's feasible: a micro-benchmark needs ≈ 103 to 106 timed calls to beat the 1 µs clock resolution; a body of ~100 steps at 106 calls = 108 steps ≈ 1 s, so size iterations to that budget.

Quiz

Interview questions

Cheat sheet

ConceptSyntax / rule
Test pyramidMany unit tests (fast, precise) → some widget tests → few integration tests (slow, end-to-end)
test / grouptest('desc', () { ... }); registers one test; group('name', () { ...tests... }) labels/nests several
Matchersequals(x), isA<T>(), throwsA(m), closeTo(x, delta), contains(x)
setUp / tearDownrun before / after EVERY test in scope — build fresh state here, never share a mutable object across tests
Async testmake the test body async and await what you're testing — a missing await reports green with no real check having happened yet
TDD loopred (failing test) → green (minimal fix) → refactor (clean up, safety net already in place)
Stub / Fake / MockStub = canned answer. Fake = real, simplified working implementation. Mock = framework double (mockito/mocktail) that also verifies HOW it was called.
Dependency injectionpass a Clock/HttpClient/etc. IN, instead of constructing it inside the function — the only way to make time/network/randomness testable
AAAArrange (build inputs/fakes) → Act (call the one thing under test) → Assert (check the outcome)
Edge casesalways test: empty, one, many, boundary, invalid/hostile input
Property-based testcheck a general property ("matches a trusted reference") against MANY random inputs, not a few hand-picked ones
Coveragedart test --coverage=coverage — measures which LINES ran, not which behaviors were actually CHECKED; 100% ≠ bug-free
Stack trace#0 = throw site (innermost); read top to bottom, each higher number is one caller further out
Step over/into/outover = stay in this function, skip past a call; into = go inside the call; out = finish this function, land in its caller
assert vs debugger()assert = debug-only internal check, stripped in release. debugger() (dart:developer) = programmatic breakpoint for an already-attached debugger.
git bisectbinary search over commit history: git bisect start, bisect bad, bisect good <c>, or automate with bisect run <script>
Flame chartbar width = time; nested bars = callees; look at SELF time (own code only), not just total width, to find the real hotspot
Stopwatch benchmarkingwarm up before timing (JIT), and feed every result into a used "sink" (avoid dead-code elimination)