Moli treats browser layout as a disposable snapshot for AI agents
Overview
My reading of the moli README is that its central bet concerns when a browser does visual work. It ships a full rendering stack, yet by default it runs no real layout or paint at all. Geometry comes back mocked unless you pass --layout, and even with that flag on, layout is a snapshot the browser builds, freezes, and holds only until the next rebuild. If your agent reads pages instead of looking at them, that default is the part to evaluate.
Moli comes from Lexmount, is written in Rust, and the README describes it as a standalone browser kernel rather than a Chromium wrapper. It lists its building blocks:
- libcurl for network transport
- html5ever for HTML parsing
- rusty_v8 / V8 for JavaScript
- Servo/Stylo for selectors, cascade, and computed style
- Taffy and Parley for box and text layout
- AnyRender/Vello CPU with usvg for software rendering
Structure by default, pixels by request
The README argues that browser automation tasks need page structure far more often than a continuously rendered visual world. Moli treats the native DOM and style state as the single source of truth and triggers layout or software paint only for operations that require them.
Its request table breaks the cost down like this:
- Extracting HTML or Markdown, querying the DOM, running JS, or inspecting network and storage - reads runtime state directly, with no layout or paint.
- Reading an element's box, hit-testing coordinates, or sending coordinate input - runs one layout calculation and keeps only the latest frozen layout tree.
- Capturing a screenshot - rebuilds from the current DOM and style, replaces the frozen tree, renders a fresh frame, and discards the paint state afterward.
- Polling a screencast - compares generation metadata; a clean state emits no frame, and a changed state produces one fresh frame.
Cost controls
The cost controls follow the same opt-in pattern:
- The default is
LayoutPolicy::Mock, which the README describes as deterministic geometry in a compatible format, with no real layout or paint. - Passing
--layoutswitches toLayoutPolicy::OnDemand.
Media is also opt-in:
--resourcefetches all optional visual and media resource families.- Flags such as
--imageor--fontenable one family each.
The frozen tree and its tradeoff
According to the README, the first geometry request builds a working layout tree from the current DOM and style, freezes its canonical geometry into an immutable, DOM-independent FrozenLayoutTree, and retains only that latest tree.
The architecture section adds that each real refresh then discards the working tree, style borrows, layout caches, diagnostics, and paint state. Paint results are never reused.
The sentence I would flag for anyone building on this:
"Ordinary geometry reads may reuse it even if the page has changed."
Screenshots always rebuild and replace the frozen tree, but a plain box read after a DOM mutation may be answered from the earlier snapshot. If your agent clicks by coordinates on a page that is still settling, test that path against your own flows before trusting the positions it returns.
One endpoint for three protocols
moli serve starts an automation server, and the README says the same endpoint serves CDP, WebDriver Classic, and WebDriver BiDi, which share one kernel and scheduler. No separate ChromeDriver, geckodriver, or browser installation is required.
The example connects Playwright over CDP with chromium.connectOverCDP.
moli serve --layout adds real geometry, coordinate input, and screenshot and screencast surfaces.
One-shot extraction
For one-shot extraction:
moli fetch --dump markdown --wait-until donerenders a page as Markdown.--dump semantic_tree_textreturns what the README calls a compact, model-friendly semantic tree.
Visual output needs the layout flag:
moli fetch --layout --dump screenshot- withscreenshot_fullandpdfas the other dump targets.
Reading the project's numbers
Every figure below is the project's own reported measurement.
In a mixed crawl of 192 public URLs from Chinese and international sites, a page counts only if it produces meaningful content after JavaScript runs.
| Engine | Useful pages | Median time | Median RSS |
|---|---|---|---|
| moli | 103 (53.6%) | 1.43 s | 73 MiB |
| Chrome Headless | 101 (52.6%) | 1.43 s | 773 MiB |
| Lightpanda | 85 | 0.97 s | 40 MiB |
Read plainly, the reported gap with Chrome Headless in that test sits in memory, while success rate and median time are close or identical.
A sample agent workload in the README puts moli's CDP ready time at 34.85 ms against 169.37 ms for Chromium, peak PSS at 102.46 MiB against 348.82 MiB, and 1 process with 24 threads against 11 processes with 123 threads.
The project also reports that one full run of its selected WPT tests passed 1.612 million tests.
Fit for purpose
The README names crawling, browser-use agents, retrieval pipelines, evaluation environments, and reinforcement-learning workloads as fits for this cost model.
If your workload depends on screenshots, benchmark it with --layout enabled, since that mode is where the rebuild steps described above run.
GitHub: https://github.com/lexmount/moli
Curated by Agent Palisade - practical AI for small and mid-sized businesses.
Comments
No comments yet. Start the discussion.