Performance & Testing in Flutter

By the end of this lesson you will be able to find out why a screen is janky instead of guessing — run the right build mode, read a DevTools frame chart and tell a UI-thread problem from a raster-thread problem — and then fix the usual causes: too many rebuilds, lists that build everything, oversized images and expensive effects. In the second half you will write the three kinds of Flutter tests an interviewer asks about: widget tests (finders, taps, typing, pump versus pumpAndSettle, Futures), golden tests (pixel snapshots) and integration tests (the real app on a device). Every printed number on this page comes from a real Flutter test.

How this lesson proves things. Every code box lives in a real test file, verify_flutter/test/f05_test.dart, run with flutter test. A widget test builds widgets on a pretend screen with no window and a fake clock (section 5 explains both). The // => lines under each box are what the test really printed. Two things cannot be measured inside a test and are stated from Flutter's documentation instead, clearly marked: real frame times on a phone (sections 1 and 4 — a test has no GPU and no real screen), and integration tests (section 7 — they need a package this project does not add). This lesson builds on Step 12.1 (rebuilds and const), Step 12.3 (layout, painting, RepaintBoundary), Step 12.4 (jank and isolates) and the Dart side track's Step S.5 (test, expect, matchers and the test pyramid in plain Dart).

1. Measure first: build modes and DevTools

A doctor does not hand out medicine because a patient "feels slow". She takes the temperature, runs one test, changes one thing, and measures again. And she measures the patient at rest, not while they are wearing a heavy backpack full of instruments — otherwise every number is wrong. A Flutter app in debug mode is wearing that backpack.

Flutter can build your app in three build modes (ways of compiling and packing the same code):

Your code can ask which mode it is in with three constants, exactly one of which is true. flutter test runs in debug mode:

Flutter DevTools is the browser-based tool kit that connects to a running app (open it from your IDE or from the link flutter run prints). Its Performance view draws a frame chart: one pair of bars per frame. The first bar is the UI thread — the main isolate from Step 12.4, running your Dart code plus build, layout and paint. The second is the raster thread, which turns the frame's layers into real pixels on the GPU (Step 12.3). A frame is janky when either bar is longer than the frame budget (16.7 ms at 60 Hz, 8.3 ms at 120 Hz); DevTools colours those frames so they stand out. Which bar is too long tells you where to look:

Measuring the wrong thing. (1) Timing in debug mode — the most common mistake; always use profile mode on a real device. (2) Changing three things at once and not knowing which one helped. (3) Optimising code that DevTools never showed as slow: a rebuild of a few small widgets costs microseconds. The loop is always the same: measure → find the longest bar → change one thing → measure again.
Where DevTools gets its numbers. In profile mode the engine records how long each frame spent on each thread and sends it to DevTools through the VM service (the debugging connection to the running app); your app can read the same numbers with SchedulerBinding.instance.addTimingsCallback, which hands you a list of FrameTiming objects. DevTools can also count rebuilds: its rebuild statistics are switched on through a debug hook called debugOnRebuildDirtyWidget — the Flutter source shows the inspector installing exactly that hook — and question q27 uses the same hook inside a test. A UI-thread problem is opened up further with the CPU profiler (which functions took the time); a raster-thread problem with the layer and shader information in the Performance view.
Measure in profile mode on a real device, never in debug. Frame chart: UI bar long → your Dart code, build, layout or paint; raster bar long → too much drawing work. Change one thing, measure again.

2. Rebuild control

A shop updates the "items in cart" number on its sign. A careless shop reprints the whole sign — logo, product list, opening hours — every time someone adds an item. A smart shop has a small separate counter display and only changes that. Flutter rebuilds are the same: the trick is to make the part that changes small, and to tell Flutter which parts are finished (const).

Remember the rule from Step 12.1: setState marks one element dirty; in the next frame its build runs, and every child widget that is not the very same object as last time is updated and builds too, all the way down. A const child returned from the same line is the same object every time, so Flutter skips it and everything below it. The test below counts our own widgets' build calls after one tap on "cart" in four versions of the same shop screen (a header, the cart counter, a list of 5 products, a footer):

The four tools, in the order to try them:

To see what rebuilds in your own app, DevTools shows rebuild counts per widget; in a debug build you can also switch on a log of every rebuilt widget:

Rebuilds are not always the problem. A build method of a small widget costs microseconds. Do not turn a readable screen into twenty tiny widgets because "rebuilds are bad" — fix the subtrees that DevTools shows rebuilding often (every frame of an animation, every keystroke) and that are big (long lists, charts). Also: const only helps when the value can really be constant; ProductTile(n) with a changing n cannot be.
Why the counts are exactly 9, 1, 1, 1. Version 1: Shop is dirty → its build makes brand-new Header(), ProductList() and Footer() objects → none is identical to the old one → each builds, and ProductList's build makes five new tiles → 1 + 1 + 1 + 5 + 1 = 9. Version 2: the const children are canonical (one shared object) → identical → skipped; only Shop builds (plus the framework's own Text, which we do not count). Version 3: Shop is not dirty at all; only CartButton's element is. Version 4: ValueListenableBuilder calls setState on itself when the notifier changes, so only its builder runs. In profile and release builds the same rule holds; debug builds only add a hidden source-location note to const widgets for the inspector (Step 12.1), which does not change identity at a single place in the code.
Fewer rebuilds: const → split into widgets / push state down → builders → builders' child parameter. Fix what DevTools shows rebuilding often and big — not everything.

3. Long lists and images

A theatre with 10 000 seats but a stage only 10 metres wide. You do not dress all 10 000 actors at once — you dress the ones on stage, plus a few waiting in the wings on each side so they are ready the moment the scene scrolls on. When an actor walks off far enough, they get changed back and their costume is reused. The wings are the cache extent.

A viewport is the window a scrolling list shows (here: a box 800 pixels tall). A list is built of slivers (pieces of a scrollable area that know how to lay out only part of themselves). ListView builds and lays out only the items that touch the visible area plus the cache extent on each side — 250 pixels by default. The difference between the two constructors is who makes the widget objects:

Try it: which items are built?

Type a list length, the viewport height, the item height, the cache extent and a scroll offset. The player runs a JavaScript copy of the Dart function builtWindow from interview question q17: an item is built when it touches the region from offset − cache to offset + viewport + cache. The test file checks that function against a real ListView.builder on every suggestion below and on 60 random inputs, counting the items that really exist after a jumpTo.

itemExtent and prototypeItem. If every item has the same height, say so: itemExtent: 80, or prototypeItem: ListItem(0) (Flutter measures that one example item once). With a known height, the list can compute where item 5 000 is — 5 000 × 80 = 400 000 pixels — without building anything above it. Without it, the list cannot know how tall the items before 5 000 are, so a jump makes it build, measure and throw away every item on the way:

Images are the other classic list problem. A photo file is compressed; to draw it, it must be decoded into raw pixels, about 4 bytes per pixel — and Flutter decodes at the image's full size unless told otherwise. A 12-megapixel photo shown as a 100-pixel thumbnail still costs 48 MB of memory and the decoding time. cacheWidth/cacheHeight tell Flutter to decode at the size you actually show (in physical pixels: logical size × device pixel ratio, question q11):

Three list traps. (1) shrinkWrap: true inside another scroll view: the list must lay out every item to know its own height (Step 12.3 measured it; question q19 shows it is fine when the list has a bounded height). Prefer one CustomScrollView with slivers. (2) Keeping state inside list items: an item that scrolls out of the cache is destroyed, and its State with it — keep the data outside the item, or use AutomaticKeepAliveClientMixin (question q20). (3) Items that can move (reorder, insert at the top) need keys so state follows the right item (Step 12.1).
Exactly which items? With a fixed itemExtent, the list works out the first index as floor((offset − cache) ÷ extent) and the last as ceil((offset + viewport + cache) ÷ extent) − 1, clamped to the list; the cache region never starts above item 0 or ends past the last item. The test found one small difference without itemExtent: an item whose bottom edge sits exactly on the start of the region is also kept (49..59 instead of 50..59 in the test). Items that are built but off screen are in the tree and laid out, but not painted on screen; that is why find.byType(ListItem) needs skipOffstage: false to count them. Since Flutter 3.41 the cache is set with scrollCacheExtent: ScrollCacheExtent.pixels(250) (or .viewport(0.5), a fraction of the viewport); the older cacheExtent: 250 is deprecated, and the analyzer in this project flags it.
Long lists: ListView.builder (creates only what is needed), itemExtent or prototypeItem when heights are equal (cheap jumps), no shrinkWrap inside another scroll view, keys for movable items. Images: cacheWidth/cacheHeight = shown size × pixel ratio.

4. Raster-side costs: layers, saveLayer and shaders

Painting a window with a semi-transparent blue tint. The quick way: mix the tint into the paint and paint once. The slow way: paint the whole scene on a separate sheet of glass first, then hold the sheet up and fade it — an extra sheet, and painting everything twice. Some widgets need that extra sheet. On the GPU it is called saveLayer: draw into an offscreen buffer, then blend the buffer onto the screen.

When the raster bar is the long one, the UI thread is fine but the frame is too expensive to draw. The paint step produces a tree of layers (Step 12.3); a widget test can list them with tester.layers. Some effects add their own layer, and the engine may need an offscreen buffer to draw that layer correctly:

What the test shows, and what Flutter's performance guide says about it:

Shader compilation jank. A shader is a small program the GPU runs to draw pixels. With Flutter's older renderer, Skia, shaders were compiled the first time an effect appeared on screen, which could stall a frame for tens of milliseconds the first time an animation ran — "first-run jank" that does not show up on the second run. Flutter's documentation describes Impeller, the newer renderer, as compiling all its shaders ahead of time, when the engine is built, so this kind of jank goes away. Impeller is the default renderer on iOS and on modern Android devices; check the current "Impeller rendering engine" page for which devices still fall back to the older renderer, and measure on those too. (None of this can be tested in this project: a widget test has no GPU.)
Why can a layer be expensive at all? The raster thread combines all layers into the final image every frame. A plain picture is drawn once into the frame. An effect that needs an offscreen buffer forces the GPU to switch to a separate texture, draw into it, switch back and blend — on mobile GPUs that work in tiles, each such switch is costly. The DevTools Performance view can highlight frames with many offscreen layers, and the debug flag debugDisableOpacityLayers (with friends such as debugDisableClipLayers) lets you switch an effect off temporarily to see whether the raster time drops — measure, change one thing, measure again.
Raster bar long → look for effects that need offscreen buffers: Opacity on groups, ShaderMask, filters, blurs, antiAliasWithSaveLayer clips. Use colour alpha or FadeTransition, keep effects small, RepaintBoundary for frequent small repaints. Shader jank: Impeller compiles shaders ahead of time.

5. Widget tests

A flight simulator. The pilot uses real controls and sees a real-looking cockpit, but nothing actually flies: there is no sky, no engine noise, and the instructor can jump the clock forward ("it is now two hours later") with one button. A widget test is a flight simulator for your widgets: real widgets, real layout, a pretend screen and a clock you control.

From Step S.5 you know test(), expect() and matchers. package:flutter_test adds:

The middle line is the most important lesson about widget tests: an action does not redraw the screen. tap runs the onPressed handler, setState marks the widget dirty — and nothing else happens until you pump a frame. That is where the fake clock comes in:

Testing a widget that uses a Future. It depends on what the Future waits for. A Future driven by a timer (Future.delayed, a debounce, a fake repository that waits 2 seconds) lives on the fake clock: pump(duration) completes it instantly. A Future that waits for the real world — reading a file, a real HTTP call, decoding an image — is finished by the operating system in real time, and the fake clock cannot speed it up; wrap the real wait in tester.runAsync(...), then pump. In practice you replace the real network with a fake (Step S.5, test doubles) so the test stays fast and certain:

Four classic widget-test bugs. (1) Forgetting pump() after a tap and wondering why the text did not change. (2) pumpAndSettle() on a screen with a never-ending animation (a CircularProgressIndicator, a shimmer, a repeating controller): it times out after 10 minutes of fake time with "pumpAndSettle timed out" — use pump(const Duration(milliseconds: 500)) or a helper like question q21's pumpUntilFound. (3) A widget with no MaterialApp/Directionality above it: Text and TextField need one (question q6). (4) Tapping something that is off screen or covered: the tap lands elsewhere and Flutter prints a warning; scroll to it first (tester.ensureVisible).
Why a fake clock? testWidgets runs your test body inside a special zone (Step S.1) whose timers belong to a fake clock. That makes a test with a 30-second timeout take milliseconds and gives the same result on every run. The price: anything that is not a Dart timer — file and network operations, isolates, image codecs — happens outside that clock, which is why runAsync exists (Step 12.4 used it for Isolate.run). The binding also checks at the end of every test that you left no timers running, no animation ticking and no debug flag switched on — the panels above switch debugPrintRebuildDirtyWidgets back off for that reason.
Widget test = testWidgets + pumpWidget + finders + actions + expect. Action → pump() to see the result. pump(d) moves fake time; pumpAndSettle waits for finite animations only. Timer Futures: pump; real I/O: runAsync (or better, a fake).

6. Golden tests

A tailor keeps a photo of the perfect suit. After every change to the pattern, she photographs the new suit from exactly the same spot, with the same lamp, and lays the two photos on top of each other. Any difference — a crooked pocket, a wrong colour — jumps out. But if she moves the lamp or uses a different camera, everything "differs" even though the suit is the same.

A golden test (also called a snapshot or screenshot test) draws a widget, takes a PNG picture of it, and compares it pixel by pixel with a saved reference picture, the golden file. You create or refresh the golden files with flutter test --update-goldens; a normal flutter test then fails if a single pixel differs and writes the difference images into a failures/ folder so you can see what changed. This lesson's golden file is verify_flutter/test/goldens/f05_badge.png, made once on this machine and checked on every run:

The badge has no text on purpose. Under the hood the comparison is GoldenFileComparator.compareLists, which reports whether two PNGs match and what fraction of pixels differ (question q23 measures 66.67 % when the green part turns teal).

Why goldens break on another computer. Text and anti-aliased edges are drawn slightly differently on macOS, Linux and Windows (different font rendering, different graphics libraries), so a golden made on one developer's Mac can fail on a Linux CI machine with no real change. Fixes that teams use: (1) generate and check goldens on one consistent machine — usually the CI server, in a fixed container image; (2) keep each golden small — one widget in one state, not a whole screen — so a failure points at one thing and is cheap to compare; (3) in tests, text uses the special test font by default, in which every letter is a square box (question q31), so real fonts do not leak in unless you load them; (4) if you must, allow a tiny tolerance with a custom comparator (question q32).
Where the comparison happens. matchesGoldenFile finds the nearest RepaintBoundary around the widget, renders its layer to an image, encodes a PNG and hands it to the global goldenFileComparator. In flutter test that is a LocalFileComparator: it treats the name as a path relative to the test file, and with --update-goldens it writes the file instead of comparing. Replacing goldenFileComparator (often in a flutter_test_config.dart file next to the tests) is how teams plug in tolerances or cloud comparison services.
expectLater(find..., matchesGoldenFile('goldens/x.png')); create with --update-goldens; small widgets; one consistent machine; expect platform and font differences.

7. Integration tests and the test pyramid

Engine parts are tested on a bench (unit tests), the assembled engine is tested on a stand (widget tests), and finally a test driver takes the whole car out on a real road (integration tests). The road test finds problems no bench can — but it is slow and expensive, so you do fewer of them.

An integration test runs your whole app on a real device or emulator: the real engine, real GPU, real platform code (plugins, permissions, the camera), real time. It catches what a widget test cannot — a plugin whose native half is missing, a permission dialog, a crash only on Android. A widget test has no native side at all (question q29 shows a plugin call failing in a widget test).

The flow itself is written exactly like a widget test. Here is a checkout journey, run as a widget test in this lesson's test file:

To run the same steps as an integration test, Flutter's documentation describes this setup (shown here, not run in this lesson: it needs the integration_test package, which ships with the Flutter SDK but must be added to pubspec.yaml):

  1. Add it as a dev dependency: under dev_dependencies: write integration_test: with sdk: flutter below it.
  2. Create the folder integration_test/ next to lib/ and test/, with a file such as checkout_test.dart.
  3. Run it on a connected device or emulator with flutter test integration_test/checkout_test.dart (or flutter test integration_test for all of them). On CI, cloud device farms such as Firebase Test Lab run the same files on many real phones.

The only differences from the widget test: the binding is IntegrationTestWidgetsFlutterBinding (it connects the test to the real app on the device), and the test starts the real app with app.main() instead of pumping one widget.

Unit testWidget testIntegration test
TestsOne function or class (plain Dart)One widget or screen on a pretend screenThe whole app on a device
SpeedMillisecondsMilliseconds to a secondSeconds to minutes (build + install + run)
Real platform code, GPU, timeNoNo (fake clock, mocked channels)Yes
How manyManyManyA few key journeys (sign-in, checkout)
Run withflutter test / dart testflutter testflutter test integration_test/ on a device
An upside-down pyramid. Teams that test everything through integration tests get a suite that takes an hour and fails randomly (a slow network, an animation that was not finished). Keep the pyramid: many fast unit and widget tests for logic and screens, a handful of integration tests for the journeys that make money. And do not use integration tests to "test performance" casually — measuring needs profile mode and repeated runs.
Integration tests can also record performance: the integration_test binding offers traceAction and watchPerformance, which collect frame timings while a scripted action runs (for example, scrolling a long list), so a CI job can fail when frames get slower. Flutter's documentation covers this under "performance profiling" for integration tests; run such tests in profile mode on a real device for meaningful numbers.
Unit (logic) → widget (one screen, fake screen and clock) → integration (whole app, real device). Many, many, few. Integration tests: integration_test dev dependency, integration_test/ folder, IntegrationTestWidgetsFlutterBinding.ensureInitialized(), flutter test integration_test/.

8. How to answer the three classic interview questions

A good answer is a short story with a symptom, a measurement, a cause and a proof that the fix worked — like a mechanic who says "the car shakes at 80 km/h; I measured the wheels; the front left is unbalanced; after balancing, no shaking at 80".
QuestionA strong answer, in order
"How do you find out why a screen is janky?"(1) Reproduce it in profile mode on a real (older) device — debug numbers lie. (2) Open DevTools' Performance view and find the janky frames. (3) Look at which bar is over budget. UI thread → CPU profiler and rebuild counts: too many rebuilds (fix with const, splitting, builders, child), a list that builds everything (ListView.builder, itemExtent), heavy work on the main isolate (Isolate.run, Step 12.4). Raster thread → expensive effects (Opacity on groups, blurs, saveLayer clips), huge images (cacheWidth), first-run shader compilation on the older renderer. (4) Change one thing and measure again.
"Unit vs widget vs integration tests?"Unit: one function or class, plain Dart, milliseconds. Widget: one widget or screen with testWidgets, a pretend screen and a fake clock — finders, tap, enterText, pump. Integration: the whole app on a device with the integration_test package — real plugins, real GPU, slow. Pyramid: many unit and widget tests, a few integration tests for key journeys; goldens for visual regressions on one consistent machine.
"How do you test a widget that uses a Future?"(1) Inject the data source so the test can pass a fake (no real network). (2) pumpWidget → expect the loading state. (3) For a timer-based Future, pump(duration) moves the fake clock; for an already-finished Future, pump(Duration.zero) delivers it. (4) Expect the data state; test the error state with a fake that throws. (5) Real I/O completes outside the fake clock — wrap it in tester.runAsync. Avoid pumpAndSettle when a spinner keeps animating.
Weak answers interviewers hear every day: "I add const everywhere" (without measuring), "I test performance in debug mode", "I use pumpAndSettle everywhere" (it hangs on spinners), "goldens are flaky so we deleted them" (run them on one machine), and "integration tests replace widget tests" (they are slow and fewer).
Follow-ups to expect: "What does RepaintBoundary do?" (Step 12.3), "Why is ListView(children:) bad for 10 000 items if Flutter only builds the visible ones?" (the 10 000 widget objects are still created up front), "How do you test code that uses a platform channel?" (setMockMethodCallHandler, Step 12.4), "How would you catch performance regressions in CI?" (integration tests with frame timing on a real device in profile mode).
Symptom → measurement (mode, DevTools, which thread) → cause → one fix → a number that proves it.

Quiz

Interview questions

Cheat sheet

ConceptFact
Build modesDebug: JIT, asserts, hot reload, slow. Profile: AOT + tracing, for measuring, real device only. Release: AOT, what users get. kDebugMode / kProfileMode / kReleaseMode.
Frame chartPer frame: UI bar (Dart, build, layout, paint) and raster bar (GPU drawing). Janky if either is over budget (16.7 ms at 60 Hz, 8.3 ms at 120 Hz).
Rebuild controlconst → split widgets / push state down → builders → builders' child. Helper methods do not stop rebuilds.
Rebuild toolsDevTools rebuild counts; debugPrintRebuildDirtyWidgets; debugOnRebuildDirtyWidget (debug only, switch off again).
ListsListView.builder creates only visible + cache items (cache 250 px each side by default, scrollCacheExtent). ListView(children:) creates all widget objects up front.
itemExtent / prototypeItemEqual heights known → jumps build only the new window (18 items, not 5 000).
shrinkWrapUnbounded (inside another scroll view): lays out every item. Prefer slivers.
ImagesDecoded ≈ width × height × 4 bytes. cacheWidth = shown width × device pixel ratio.
Raster costsOpacity on groups, ShaderMask, ColorFiltered, BackdropFilter, ImageFiltered, Clip.antiAliasWithSaveLayer. Use colour alpha, FadeTransition. Impeller: shaders compiled ahead of time.
Widget testtestWidgets, pumpWidget, find.text/byKey/byType, tap, enterText, findsOneWidget. Action → pump().
Pumpingpump() one frame; pump(d) moves fake time; pumpAndSettle 100 ms steps until idle, times out on endless animations.
Futures in testsTimers: pump(d); finished Future: pump(Duration.zero); real I/O: tester.runAsync; best: inject a fake.
Golden testsmatchesGoldenFile; --update-goldens; small; one consistent machine; text is boxes in the test font.
Integration testsintegration_test dev dependency; integration_test/ folder; IntegrationTestWidgetsFlutterBinding.ensureInitialized(); flutter test integration_test/ on a device.