The app worked perfectly on every emulator, every office Wi-Fi network, and every staging build. In production, on cellular connections above roughly 150ms round-trip, one in every few hundred sessions froze on the home screen with a spinner that never resolved. No crash log, no exception, no ANR trace worth reading. Just a UI that stopped updating and a provider tree that, if you’d been watching it, had quietly wedged itself into a cycle it could never escape. That’s the failure mode we want to walk through, because it’s a class of bug that unit tests won’t catch and that only shows up when two async providers disagree about who goes first.

The setup: two providers, one implicit contract

The architecture was standard Riverpod: a sessionProvider (an AsyncNotifierProvider that resolves the authenticated user and refreshes the token when it’s near expiry) and a workspaceProvider (an AsyncNotifierProvider that loads the active workspace, which requires a valid session token to call the API). The dependency looked clean on paper:

  • workspaceProvider reads sessionProvider to get the token.
  • If sessionProvider is still loading, workspaceProvider awaits it via ref.watch(sessionProvider.future).
  • If the token is near expiry, sessionProvider‘s build method kicks off a refresh, invalidating itself once the refresh completes.

The bug was in what “near expiry” triggered. The refresh logic didn’t just update internal state — it called ref.invalidateSelf() partway through its own build, after having already read a value from workspaceProvider to decide whether a refresh was safe to run concurrently with an in-flight workspace fetch (a guard someone had added to avoid hammering the auth endpoint during a workspace switch). That single read created a cycle: workspaceProvider depends on sessionProvider, and sessionProvider‘s build now depended on workspaceProvider.

Riverpod doesn’t reject this at compile time — the cyclical read only exists at runtime, conditionally, inside a branch that only executes when a token is close to expiring. Under fast networks, both providers resolved before either one hit the branch that read the other, so the cycle was never actually walked. It took real latency — a session refresh call and a workspace fetch both in flight for more than a couple hundred milliseconds — for both providers to be mid-build at the same time, each awaiting the other’s future. That’s the deadlock: two providers, both suspended, waiting on each other, and no timeout anywhere in the stack to break it.

Why this hid from every normal debugging pass

The reason this survived code review and QA is that nothing about it looks wrong locally:

  • Each provider, read in isolation, has a straightforward, linear dependency.
  • The offending read (sessionProvider checking workspaceProvider‘s state) was several call-levels deep in a helper function, not visible in the same file as the ref.watch calls that defined the “official” dependency graph.
  • Riverpod’s own DevTools graph view showed the static dependency edges, but the guard condition was runtime-only — it didn’t show up as an edge because it wasn’t a ref.watch, it was a one-off ref.read(workspaceProvider) inside a conditional, which Riverpod tracks for rebuilds but doesn’t surface distinctly in the graph visualization.

We found it by building a deliberately dumb debugging tool rather than trusting the graph view: a logging wrapper around every ref.watch and ref.read call in the two suspect providers, timestamped, printing the provider name, the call site, and whether the call resolved or suspended. Running that log against a network-throttled build (Chrome DevTools-style throttling profiles, replicated on-device with a proxy that injected 150-300ms of latency) reproduced the freeze in under ten attempts. The log made the cycle obvious in a way the static graph couldn’t: sessionProvider build starts → reads workspaceProvider → suspends → workspaceProvider build starts → reads sessionProvider.future → suspends → nothing left to drive either future forward.

The fix, and the actual lesson

The fix itself was small: move the “don’t refresh mid-fetch” guard out of the provider graph entirely and into a plain in-memory flag set directly by the workspace fetch method, read via ref.read outside of any provider’s build phase. That breaks the cycle because the guard no longer participates in Riverpod’s dependency tracking at all — it’s just shared mutable state checked at the right moment, which is honestly what it always should have been.

The broader lesson is about where cross-provider coordination logic lives. Riverpod makes it very easy to reach into another provider’s state from anywhere, and that flexibility is exactly what let a guard condition smuggle a hidden dependency into the graph. Any time a provider’s build method reads another provider conditionally — not as its primary data dependency, but as a side-check — that’s worth treating as a design smell, because it’s invisible to every tool that assumes the graph is static. We now treat conditional cross-provider reads inside build methods as something to flag in review specifically, independent of whether the code “works” in testing, because working in testing is exactly what this bug did for months.

This is the kind of failure that doesn’t show up in a demo and doesn’t show up in a code sample — it shows up in production, under real network conditions, in a codebase that already has history. It’s also the kind of thing we spend a lot of our audit time on at orithLabs: not style nits, but tracing the actual runtime shape of state management against what the code appears to promise. If you’ve got a Riverpod, Bloc, or Redux-based Flutter app with intermittent freezes that never reproduce on demand, that mismatch between apparent and actual dependency graphs is usually where to start looking.