The bug report said “sometimes the cart total is wrong after rotating the screen.” Not a crash, not a stack trace — just a number that occasionally didn’t match. It took four days to reproduce, and it only reproduced on one device class: Samsung Fold-series phones, and only when the app was open across a fold-state transition. Every unit test passed. Every widget test passed. Golden tests passed. CI was green the entire time the bug was live in production.

This is a common shape of failure in Flutter apps that lean on platform channels for anything beyond trivial data, and it’s worth walking through why the standard test pyramid misses it entirely.

Where the corruption actually happened

The app used a MethodChannel to hand off camera capture and image metadata to native code (Kotlin on Android, Swift on iOS), because the Flutter-side camera plugin didn’t expose EXIF orientation data reliably. The native side packaged a result as a Map<String, Any> and sent it back over the channel using standard codec serialization.

On a fold-state change, Android can reconfigure the Activity — depending on manifest configuration, a fold/unfold can trigger a full onConfigurationChanged or, in some OEM skins, an actual Activity recreation. The FlutterEngine can survive this, but the native-side singleton holding in-flight capture state doesn’t, unless you’ve deliberately wired it to survive. On Samsung’s Fold devices specifically, the recreation window overlapped with an async native callback still in flight from before the fold event. That callback wrote into a state object that had already been replaced by the newly recreated Activity’s version.

The result: the MethodChannel call returned successfully, with a well-formed, correctly typed map. Nothing about the returned value was invalid. It just described the state of the previous Activity instance, not the current one. On the Dart side, this looked like ordinary data — a valid image path, a valid timestamp — that happened to belong to the wrong capture session.

Why the test suite couldn’t see it

Standard Flutter test layers are structurally blind to this class of bug:

  • Unit tests mock the MethodChannel via TestDefaultBinaryMessengerBinding or a fake handler. They verify that Dart code reacts correctly to a given payload shape — they never touch the actual native serialization path, let alone Activity lifecycle timing.
  • Widget tests run in a Dart VM with no real Android or iOS host. There’s no Activity to recreate, no configuration change to race against.
  • Integration tests (integration_test package) run on a real device, which is closer, but almost nobody writes an integration test that specifically triggers a fold/unfold mid-async-call, because it requires either physical foldable hardware or an emulator fold-state script, and most CI fleets don’t have either.

The failure mode wasn’t a logic error. It was a lifecycle race that only exists at the native boundary, on specific hardware, under a specific interaction sequence. No amount of additional Dart-side test coverage would have caught it, because the bug wasn’t in the Dart-side code at all — it was in the assumption that a MethodChannel result is always scoped to the call that requested it.

What actually catches this class of bug

The fix had two parts, and both matter more than the specific bug: the native-side state was moved out of an Activity-scoped singleton into a ViewModel-scoped holder that survives configuration changes, and every MethodChannel call that returns async results got a request-id round trip — the Dart side tags the call, native echoes the tag back, and any result with a stale tag is dropped rather than applied.

The tagging pattern is the part worth generalizing. Any MethodChannel call where the native side does meaningful async work — camera, biometrics, background location, native crypto operations — should treat “the call returned” and “the call is still relevant” as separate questions. This isn’t a Dart problem or a Kotlin problem in isolation; it’s a boundary-ownership problem, and it needs to be tested as one.

Practically, that means:

  1. Writing native-side instrumented tests (Espresso / XCTest) that explicitly simulate configuration change or process death mid-callback, not just Dart-side mocked-channel tests.
  2. Treating foldables, split-screen, and multi-window as first-class configuration-change triggers in test plans, not edge cases — they trigger the same lifecycle events as rotation, just less predictably.
  3. Auditing any platform channel that carries mutable session state for implicit assumptions about call/response affinity, since the codec will happily serialize a perfectly valid answer to the wrong question.

This is the kind of bug that a code audit is actually good at catching before it ships, because it’s found by reading the shape of the native boundary code and asking “what happens if this Activity dies mid-flight,” not by running the existing test suite harder.

We’ve hit versions of this exact race in our own Flutter work — Crumb Count does camera-driven capture with native metadata handoff, and Kompete has real-time state that has to survive OEM-specific lifecycle quirks. If you’re seeing intermittent, device-specific state bugs that your CI can’t reproduce, that’s usually a native-boundary problem wearing a Dart-shaped disguise, and it’s the kind of thing orithLabs digs into directly with the people who wrote the platform channel code, not through a generic bug-triage queue.