I built a tool that finds the minimal cause of a browser bug by replaying it
Overview
Causality Cage is a tool that automates the process of identifying the minimal cause of a browser bug by replaying the failing execution flow. Instead of manually commenting out variables until something changes, you provide both a failing session and a passing session, and the tool replays the flow in Chromium with differences between the two sessions switched on and off. It then uses hierarchical delta debugging to shrink the candidate set, and iteratively re-checks by removing each component and replaying to find the root cause.
How It Works
The tool operates through several key stages:
- Session comparison - You supply a failing session and a passing session.
- Differential replay - It replays the Chromium flow with the differences between the sessions toggled on and off, covering API response fields, latency, response order, localStorage, cookies, and environment.
- Hierarchical delta debugging - This technique shrinks the set of potential causes step by step.
- Iterative verification - After narrowing down candidates, it removes each piece individually and re-plays to confirm whether the removal actually eliminates the failure.
The final report is a single offline HTML file containing a causal graph. Each factor can be clicked to neutralize, causing the corresponding failure node to flip from red to green if the change resolves the issue.
Example Case
A concrete demonstration was run on a closed issue in the RealWorld React app: gothinkster/react-redux-realworld-example-app#187. The application crashes with the error "Cannot read property 'tags' of undefined", and the reporter traced the problem to a stale token in localStorage. Because the original API host was unavailable, the author redirected the app to a public RealWorld API. On the issue's commit, Causality Cage identified the full chain of contributing factors in just 19 experiments (25.1 seconds): the stale token and the API rejecting the two requests it triggers.
An earlier version of the tool incorrectly reported only "the token" as the minimal cause, but its minimal configuration produced a different error. The current version now fingerprints the original failure and only counts runs that fail in the same manner. A flag called --loose-oracle restores the older behavior, which correctly identifies the token alone in 10 experiments.
During testing, the same run also revealed a redaction leak in the author's own report, which has since been fixed.
Results and Benchmarks
On a benchmark of planted bugs in the author's demo app, the tool achieved perfect accuracy on 10 of 10 rows, though with important caveats:
- OR row (two causes that crash with different messages) requires the
--loose-oracleflag to pass. - Flaky-oracle row initially failed but passed after adding a failure-matching rule that excludes runs which crash before reaching the target failure from the fail rate calculation.
- Latency race flip point was observed at 808 ms against an expected range of 600-900 ms.
The author notes that the app is their own implementation, so these metrics should be treated as a self-check rather than definitive evidence of generalizability across other projects.
Limitations
It is important to understand the boundaries of Causality Cage's capabilities:
- A result is necessary and sufficient only within the specific factors it models, under your particular flow and oracle, at the measured repeat rates.
- The solution is 1-minimal, meaning it finds the smallest subset of factors that explains the failure, but it is not guaranteed to be globally smallest.
- It does not model server-side state, WebSockets, IndexedDB, or static assets.
- When the true cause lies outside the modeled scope, the tool outputs "unexplained" and includes a scope box in every report.
Installation and Usage
The tool can be installed globally with:
npm install -g causality-cage
Chromium must be available via:
npx playwright install chromium
To run the tool directly:
cage doctor
The raw output is available in the repository at launch/real-world.md. For further investigation, a standalone Playwright test can be exported that fails until the underlying bug is resolved.
Repository: https://github.com/zaydmulani09/causality-cage (MIT license)
Comments
No comments yet. Start the discussion.