Moli treats browser layout as a disposable snapshot for AI agents
DEV Community

Moli treats browser layout as a disposable snapshot for AI agents

Overview

My reading of the moli README is that its central bet concerns when a browser does visual work. It ships a full rendering stack, yet by default it runs no real layout or paint at all. Geometry comes back mocked unless you pass --layout, and even with that flag on, layout is a snapshot the browser builds, freezes, and holds only until the next rebuild. If your agent reads pages instead of looking at them, that default is the part to evaluate.

Moli comes from Lexmount, is written in Rust, and the README describes it as a standalone browser kernel rather than a Chromium wrapper. It lists its building blocks:

  • libcurl for network transport
  • html5ever for HTML parsing
  • rusty_v8 / V8 for JavaScript
  • Servo/Stylo for selectors, cascade, and computed style
  • Taffy and Parley for box and text layout
  • AnyRender/Vello CPU with usvg for software rendering

Structure by default, pixels by request

The README argues that browser automation tasks need page structure far more often than a continuously rendered visual world. Moli treats the native DOM and style state as the single source of truth and triggers layout or software paint only for operations that require them.

Its request table breaks the cost down like this:

  • Extracting HTML or Markdown, querying the DOM, running JS, or inspecting network and storage - reads runtime state directly, with no layout or paint.
  • Reading an element's box, hit-testing coordinates, or sending coordinate input - runs one layout calculation and keeps only the latest frozen layout tree.
  • Capturing a screenshot - rebuilds from the current DOM and style, replaces the frozen tree, renders a fresh frame, and discards the paint state afterward.
  • Polling a screencast - compares generation metadata; a clean state emits no frame, and a changed state produces one fresh frame.

Cost controls

The cost controls follow the same opt-in pattern:

  • The default is LayoutPolicy::Mock, which the README describes as deterministic geometry in a compatible format, with no real layout or paint.
  • Passing --layout switches to LayoutPolicy::OnDemand.

Media is also opt-in:

  • --resource fetches all optional visual and media resource families.
  • Flags such as --image or --font enable one family each.

The frozen tree and its tradeoff

According to the README, the first geometry request builds a working layout tree from the current DOM and style, freezes its canonical geometry into an immutable, DOM-independent FrozenLayoutTree, and retains only that latest tree.

The architecture section adds that each real refresh then discards the working tree, style borrows, layout caches, diagnostics, and paint state. Paint results are never reused.

The sentence I would flag for anyone building on this:

"Ordinary geometry reads may reuse it even if the page has changed."

Screenshots always rebuild and replace the frozen tree, but a plain box read after a DOM mutation may be answered from the earlier snapshot. If your agent clicks by coordinates on a page that is still settling, test that path against your own flows before trusting the positions it returns.

One endpoint for three protocols

moli serve starts an automation server, and the README says the same endpoint serves CDP, WebDriver Classic, and WebDriver BiDi, which share one kernel and scheduler. No separate ChromeDriver, geckodriver, or browser installation is required.

The example connects Playwright over CDP with chromium.connectOverCDP.

moli serve --layout adds real geometry, coordinate input, and screenshot and screencast surfaces.

One-shot extraction

For one-shot extraction:

  • moli fetch --dump markdown --wait-until done renders a page as Markdown.
  • --dump semantic_tree_text returns what the README calls a compact, model-friendly semantic tree.

Visual output needs the layout flag:

  • moli fetch --layout --dump screenshot - with screenshot_full and pdf as the other dump targets.

Reading the project's numbers

Every figure below is the project's own reported measurement.

In a mixed crawl of 192 public URLs from Chinese and international sites, a page counts only if it produces meaningful content after JavaScript runs.

Engine Useful pages Median time Median RSS
moli 103 (53.6%) 1.43 s 73 MiB
Chrome Headless 101 (52.6%) 1.43 s 773 MiB
Lightpanda 85 0.97 s 40 MiB

Read plainly, the reported gap with Chrome Headless in that test sits in memory, while success rate and median time are close or identical.

A sample agent workload in the README puts moli's CDP ready time at 34.85 ms against 169.37 ms for Chromium, peak PSS at 102.46 MiB against 348.82 MiB, and 1 process with 24 threads against 11 processes with 123 threads.

The project also reports that one full run of its selected WPT tests passed 1.612 million tests.

Fit for purpose

The README names crawling, browser-use agents, retrieval pipelines, evaluation environments, and reinforcement-learning workloads as fits for this cost model.

If your workload depends on screenshots, benchmark it with --layout enabled, since that mode is where the rebuild steps described above run.


GitHub: https://github.com/lexmount/moli

Curated by Agent Palisade - practical AI for small and mid-sized businesses.

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.