On Crumb Count, a background image-processing call on Android would occasionally just… stop returning. No exception in Dart, no crash log, no ANR. The Flutter side sat in an awaited future that never resolved, because the native coroutine backing it had thrown inside a try block that swallowed the exception and never called back to the channel. From the user’s perspective, the button just stopped working. From ours, there was nothing in the stack trace to even start debugging with. That bug cost us the better part of a day, and it’s the reason every method channel we write now follows the same contract, no exceptions.
Method channels are a seductively simple abstraction: call a method name with arguments, get a result back, like a local RPC. But the platform channel is a serialization boundary between two runtimes with completely different failure semantics, and Flutter’s docs don’t tell you what happens when the native side misbehaves instead of just erroring cleanly. We learned the gaps the hard way across Crumb Count and Kompete. Here’s the contract we now enforce on every channel we ship.
Every native call gets three explicit outcomes, never two
The naive mental model is success or PlatformException. In practice there’s a third outcome that Dart’s type system won’t warn you about: no response at all. This happens when native code deadlocks, when a coroutine/dispatch queue is cancelled mid-flight, or when an uncaught exception on a background thread kills the call without ever reaching the MethodChannel result callback. Dart’s Future just hangs.
Our contract treats these as three distinct, handled paths:
- Success — a typed, versioned payload (more on that below), never a bare Map.
- Known failure — a PlatformException with an error code from a shared enum we maintain in both Dart and native, not a free-text string we pattern-match on later.
- Timeout — every single invokeMethod call is wrapped in a hard timeout at the Dart call site. No exceptions for “this one’s usually fast.” The call that hung on us was, by every prior assumption, always fast.
The timeout wrapper is a few lines, but the discipline is in never skipping it:
Future<T> callNative<T>(String method, [dynamic args]) async {
try {
final result = await _channel
.invokeMethod(method, args)
.timeout(_channelTimeout, onTimeout: () =>
throw NativeChannelTimeout(method));
return _decode<T>(result);
} on PlatformException catch (e) {
throw NativeChannelError.fromPlatformException(e);
}
}
That NativeChannelTimeout is a real, logged, telemetry-worthy event distinct from a PlatformException. On Kompete, timeouts turned out to correlate almost perfectly with a specific Android OEM’s aggressive background-thread throttling — a pattern we’d never have spotted if timeouts were just another flavor of “it failed.”
Payloads are versioned and validated on both sides
The other failure mode is quieter: a native-side change adds a field, renames a key, or changes an int to a String, and Dart happily deserializes garbage because Map<String, dynamic> doesn’t complain until three call sites downstream. We now require every channel payload to carry an explicit schemaVersion field, checked on both ends before the rest of the payload is touched.
Concretely, each method has a corresponding Dart class with a fromChannelMap factory that throws a specific ChannelSchemaMismatch error rather than a generic type-cast failure if the version doesn’t match what the Dart build expects. This sounds like overhead for what’s often a 3-4 field response, and it is a little — but a type-cast exception three call sites away from the channel invocation is a debugging session, and a ChannelSchemaMismatch thrown at the boundary with the method name and both version numbers attached is a five-minute fix. We adopted this after a native library upgrade on Crumb Count silently changed a timestamp field from seconds to milliseconds; Dart didn’t error, it just displayed dates fifty years in the future for a subset of users until someone noticed.
Native code never fails silently, by policy
The Dart-side discipline only works if native code holds up its end. Our rule for the Kotlin/Swift side: any catch block inside a method channel handler that doesn’t rethrow to the channel’s result.error(…) is treated as a bug during code review, not a style preference. It’s tempting on the native side to log and swallow, especially when a try/catch was added defensively for a case someone thought was unreachable. But “unreachable” is exactly the case that produces the silent hang, because the Dart Future is still out there waiting for a callback that will never come.
We also standardized on structured error codes shared via a generated enum (a small codegen script keeps a single source-of-truth YAML in sync across Dart and native) rather than raw strings, because string-matching on error messages is the kind of thing that works fine until someone fixes a typo in an error message and silently breaks four call sites that were pattern-matching on it.
None of this is exotic. It’s the same rigor you’d expect at any network boundary — explicit timeouts, versioned contracts, no silent failure paths — applied to a boundary that feels local because it’s on the same device, but behaves like a distributed system underneath. That’s the pattern we keep finding at orithLabs: the bugs that actually cost teams time aren’t usually in the business logic, they’re in the boundaries nobody thought to treat with that level of care. If you’re debugging a Flutter app that “just hangs” sometimes, or auditing one before you inherit it, this is one of the first places we’d look.