Webwright ของ Microsoft: ให้ AI เขียนโค้ดท่องเว็บ แทนที่จะนั่งคลิกทีละครั้ง
DEV Community

Webwright ของ Microsoft: ให้ AI เขียนโค้ดท่องเว็บ แทนที่จะนั่งคลิกทีละครั้ง

Here's a thinking process:

  1. Analyze the Request:

    • Source: A raw article text about "Webwright of Microsoft: Let AI write code to browse the web instead of clicking one by one"
    • Task: Reformat into clean, readable Markdown for a technical audience
    • Rules:
    • Break into short, well-structured paragraphs separated by blank lines
    • Add Markdown section headings (## Heading) where topic clearly shifts
    • Use bullet/numbered lists for enumerations, inline code/fenced code blocks for code/commands/identifiers
    • Preserve every fact, name, number, quotation exactly as given. Never invent, add, remove, or reorder information.
    • Only restructure and lightly polish wording for flow and grammar; no intro, conclusion, opinion, or own commentary
    • Do not repeat the article title as a heading and do not wrap the whole answer in a code fence
    • Output only the reformatted article body as Markdown, nothing else
  2. Analyze the Input Text:

    • The text is already structured but messy, with inline citations, mixed languages (Thai/English), and run-on sentences.
    • I need to parse it carefully, keep all facts/numbers/names exactly as they are, and reorganize into Markdown with headings, lists, etc.
    • Key elements to preserve:
    • Title: "Webwright ของ Microsoft: ให้ AI เขียนโค้ดท่องเว็บ แทนที่จะนั่งคลิกทีละครั้ง" (but I should NOT repeat it as a heading, per rule: "Do not repeat the article title as a heading")
    • Author/date: "โดย Nokka (นก-กา) | 3 ตุลาคม 2569"
    • Body content with many numbers, percentages, citations [1], [2], [3], [4]
    • Specific technical details: GPT-5.4, Claude Opus 4.7, Online-Mind2Web, Odysseys, AutoEval, WebArena, Skill Factory, etc.
    • Architecture: terminal, browser, model, Runner, Model Endpoint, Environment
    • Limitations/costs: token usage, dollar costs, file counts
    • Four warnings/cautions
    • Three questions to ask
    • Further reading links
  • I need to be extremely careful not to change any numbers, names, or facts. The text has some Thai phrases mixed in, which I must keep exactly as is.
  1. Identify Topic Shifts for Headings:
    The article naturally shifts between:
    • Introduction/problem statement
    • The core Webwright concept (terminal + code generation)
    • Architecture/components
    • Test results/numbers (Online-Mind2Web, Odysseys)
    • Costs and limitations
    • Architectural details/discrepancies
    • Skill Factory and reusable tools
    • Practical takeaways for Thai users
    • Questions to ask before deciding
    • Final recommendations and further reading

I'll create Markdown headings that reflect these shifts, without repeating the title. I need to make sure headings are descriptive but derived from content, not the title itself.

Let me outline potential headings (I'll refine as I go):

  • Maybe: "Problem with Current Web Agent Design"
  • "Webwright's Core Philosophy: Separating Agent from Browser"
  • "Architecture of Webwright"
  • "Test Results: Online-Mind2Web and Odysseys"
  • "Costs and Limitations"
  • "Architectural Size Discrepancies"
  • "Skill Factory and Reusable Tools"
  • "Practical Takeaways for Thai Developers"
  • "Questions to Consider Before Adoption"
  • "Further Reading"

But I must ensure headings aren't the title. I'll use short, topic-based headings.

Important: The rule says "Add Markdown section headings (## Heading) where the topic clearly shifts, so the article scans well." So I can add ## headings. I should not add an introduction or conclusion outside the content. I'll restructure the existing text into sections with headings, keeping all content verbatim (except formatting changes like bullet lists, code blocks, etc.)

Let me read the text very carefully and extract/structure it.

The text starts with: "Webwright ของ Microsoft: ให้ AI เขียนโค้ดท่องเว็บ แทนที่จะนั่งคลิกทีละครั้ง โดย Nokka (นก-กา) | 3 ตุลาคม 2569"
Then the body.

I need to preserve every fact, name, number, quotation exactly. So I'll keep [1], [2], [3], [4] citations as they appear. I'll keep all numbers, percentages, names.

Let me break the text into logical paragraphs/sections, keeping the exact wording but organizing.

Strategy: I'll copy the entire text, then rewrite it in Markdown with headings and lists, ensuring no fact is altered. I'll be very meticulous.

Let's parse the text paragraph by paragraph (mentally):

  • First paragraph: "Webwright ของ Microsoft: ... โดย Nokka (นก-กา) | 3 ตุลาคม 2569 ถ้าคุณเคยปล่อยให้ AI agent ทำงานบนเว็บยาว ๆ แล้วเห็นมันเริ่มสับสนตอนคลิกที่สี่สิบ คุณไม่ได้คิดไปเอง และงานวิจัยชิ้นหนึ่งจาก Microsoft Research อธิบายว่าปัญหาไม่ได้อยู่ที่การคลิกผิดครั้งใดครั้งหนึ่ง ปัญหาอยู่ที่ วิธีที่ agent ทำงานทั้งกระบวนการ คือดูหน้าเว็บ ตัดสินใจหนึ่งการกระทำ รอผล แล้ววนใหม่ ไปเรื่อย ๆ โดยไม่มีแผนที่ทนทานพอจะพาไปถึงปลายทาง Webwright แก้ที่รากของวิธีคิดนั้น ด้วยข้อเสนอที่ฟังดูเรียบง่ายจนน่าตกใจ คือ ให้โมเดลมี terminal แล้วให้มันเขียนโปรแกรมท่องเว็บเอง [1] ผลลัพธ์ที่ได้คือบนงานยาว ๆ โมเดล GPT-5.4 ตัวเดียวกัน กระโดดจาก 33.5% ไปเป็น 60.1% และสิ่งที่มันทิ้งไว้ให้คุณไม่ใช่รอยคลิก แต่เป็นเครื่องมือที่เอาไปรันซ้ำได้ [2] ภาพแนวคิด คือเครื่องมือที่ทำครั้งเดียวแล้วหยิบมาใช้ซ้ำได้ เทียบกับรอยคลิกที่จางหายไปพร้อมกับ session ปัญหาที่ทุกคนที่สร้าง agent เจอ แต่ไม่ค่อยมีใครตั้งชื่อให้มัน Web agent ทุกวันนี้ใช้รูปแบบเดียวกัน คือให้ browser session เป็นพื้นที่ทำงานของ agent ตัวมันเอง ในแต่ละก้าว โมเดลรับสภาพหน้าปัจจุบัน แล้วทำนายการกระทำถัดไปหนึ่งอย่าง การกระทำนั้นอาจเป็นคลิก พิมพ์ เลื่อนหน้า เลือก element ผ่าน DOM หรือเรียกเครื่องมือสั้น ๆ แต่ทั้งหมดมีข้อจำกัดร่วมกันคือ agent ถูกบังคับให้ทำนายการกระทำทีละก้าว ภายในลูปที่กำหนดไว้ล่วงหน้า [1]"

This is a big block. I'll split it into meaningful chunks, but keep wording exact. I can use bullet points for enumerations where appropriate, but must preserve all numbers/names.

Let me plan the Markdown structure:

## Problem with Current Web Agent Design
...text...

## Webwright's Core Concept: Terminal and Code Generation
...text...

## Architecture of Webwright
...text...

## Test Results: Online-Mind2Web and Odysseys
...text...

## Costs and Limitations
...text...

## Architectural Size Discrepancies
...text...

## Skill Factory and Reusable Tools
...text...

## Practical Takeaways for Thai Developers
...text...

## Questions to Consider Before Adoption
...text...

## Further Reading
...text...

But I must ensure I don't add anything outside the original text. The "Further Reading" section has URLs and such, I'll keep them as-is.

Let me carefully go through the text and reorganize, paragraph by paragraph, making sure every number, name, citation is preserved exactly. I'll also use inline code for commands/identifiers like "GPT-5.4", "Claude Opus 4.7", "Online-Mind2Web", "Odysseys", "AutoEval", "Playwright", "Chromium", "GitHub", etc. The rule says "Use inline code or fenced code blocks where code, commands, or identifiers appear." So I should format identifiers as inline code. I need to identify which parts are identifiers/commands. Things like "GPT-5.4", "Claude Opus 4.7", "Online-Mind2Web", "Odysseys", "AutoEval", "Playwright", "Chromium", "GitHub", "Microsoft", "Webwright", "Runner", "Model Endpoint", "Environment", "terminal", "browser", "skill", "CLI", "AutoEval", etc. I'll use inline code ... for those.

Also, for lists of enumerations, I'll use bullet points or numbered lists, but must preserve the exact items and order. I have to be careful not to change the meaning or order.

Let me start reconstructing. I'll write the Markdown in a text editor mentally, but I'll output directly.

First, I'll strip the title from being a heading. The rule: "Do not repeat the article title as a heading". So I won't have ## Webwright ของ Microsoft... as a heading. I'll just start with the body content, perhaps with the first heading. The title itself might appear at the top if I'm not careful, but I should not output the title as a heading. The instruction: "Output only the reformatted article body as Markdown, nothing else." So I should not include the title line at all, or if it's part of the text, I should integrate it without making it a heading. The text starts with "Webwright ของ Microsoft: ...", which includes the title. I'll treat the whole thing as the article body, and I'll not add a heading that repeats the title. I can have the first heading be something else. I might even keep the author/date as a paragraph, but not as a heading titled after the article title.

Let me re-read: "Do not repeat the article title as a heading and do not wrap the whole answer in a code fence." So I shouldn't have ## Webwright ของ Microsoft: ให้ AI เขียนโค้ดท่องเว็บ แทนที่จะนั่งคลิกทีละครั้ง. I'll make sure the first heading is not that.

I'll structure like this:

[Author and date as a paragraph, or integrated]

## Problem Statement: Limitations of Current Web Agents
...

## Webwright's Design: Separating Agent from Browser
...

## Architecture: Three Core Components
...

## Test Results: Online-Mind2Web and Odysseys
...

## Costs and Economic Considerations
...

## Architectural Footprint: File and Module Counts
...

## Skill Factory and Reusable Tooling
...

## Practical Takeaways for Repetitive Web Tasks
...

## Questions to Ask Before Adoption
...

## Further Reading
...

But I must ensure every fact from the original is preserved exactly. I'll go paragraph by paragraph, copy exactly, and just add formatting.

Let me carefully copy the text and then format. Since this is a mental exercise, I'll simulate the output.

I need to be extremely precise. Let me list out all the key info chunks to preserve:

  • Author: Nokka (นก-กา)
  • Date: 3 ตุลาคม 2569
  • Problem: AI agent gets confused after ~40 clicks on long web tasks
  • Research: Microsoft Research
  • Problem not: clicking wrong, but the loop-based design: see page, decide one action, wait, repeat, without durable map
  • Webwright solution: give model terminal, let it write web browsing code itself [1]
  • Result: GPT-5.4 alone jumps from 33.5% to 60.1% on long tasks; leaves tool reusable instead of click trail [2]
  • Concept: tool once, reuse vs click trail that vanishes with session
  • Current web agents: use browser session as agent's workspace; each step model predicts one action: click, type, scroll, select DOM element, call short tool; constrained to predict one action within predefined loop [1]
  • Researchers: this design was useful when LLMs were weak; harness helped bridge gap; but as models get better at coding, harness becomes bottleneck locking agent in narrow loop; model no longer needed in it [1]
  • When done, agent leaves only single-use click sequence, nothing reusable [3]
  • Idea: separate agent from browser; browser as tool agent can open, inspect, discard while programming [1]
  • What remains: code and records in local workspace; click finishes task once, code finishes and preserves method [3]
  • Architecture simplicity: only three things; no multi-agent, no graph engine, no plugin layer, no hidden orchestration; just terminal, browser, model [3]
  • Core system has three parts [2]:
    • Runner: stores what task is, how far, and results of previous commands
    • Model Endpoint: connects Webwright to model, supports OpenAI, Anthropic, OpenRouter
    • Environment: terminal connected to Playwright on Chromium; real command execution, real file storage, real screenshot capture
  • Short loop: understand current state, select command, run it, see what happens, repeat until model thinks done; plus self-check layer [3]
  • Numbers on Online-Mind2Web (300 real web tasks): GPT-5.4 + Webwright gets 86.7%, highest among open-source harnesses measured by AutoEval; Claude Opus 4.7 gets 84.7% but stronger on hard set 80.5% vs 76.6% of GPT-5.4 [3]
  • More important numbers: Odysseys set, 200 long tasks
    • GPT-5.4 controlling browser by screen coordinates: 33.5% [3]
    • Same model via Webwright writing code: 60.1% using avg 76.1 steps [3]
    • Increase of 26.6 points from changing harness, not model [1][3]
    • Strategic signal: better tools let smaller models work; reduces need for larger models later [3]
  • Caveats: not free; article itself lists limitations
    1. All numbers judged by LLM AutoEval, not unit tests every task; Mind2Web news uses only 100 of 300 tasks [2]
    2. Cost per task still high: ~$2.37 with GPT-5.4, $6.09 with Claude Opus 4.7; Webwright invests early compute to build durable tool, then reuses [2]
    3. Cost doesn't disappear; moves location: Microsoft using as skill on Codex uses ~3.3M tokens vs 424K when running as single harness; ~8x more because context cached in host session [2]
    4. Two points where source reports don't match:
      • System size: dev team's block says core ~1,000 lines: Runner ~150, Model Endpoint ~550, Environment ~300 [1]; README counts: main agent loop ~450 lines, Playwright env ~570, CLI 150 [3]. First counts by module, second by file; combined similar but different units. [1][3]
      • Odysseys scores: project writes 60.8%, repo/analysis uses 60.1%; analysis chose repo number for consistency; likely just rounding difference; reporting both lets you check [3]
  • Project goes further: not just finishing tasks but tool factory
    • Version 21 July 2026: Skill Factory feature; on successful problem solve, it distills script into verified, parameterized, runnable tool without model; takes ~40 seconds, no tokens used; on WebArena, reusing old tool raises accuracy from 55% to 70% (+15 points) [3]
    • Since 6 May 2026: plugin for Claude Code, Codex, OpenClaw, Hermes Agent; using single skill folder loadable across tools [3]
  • For Thai users: four most worthwhile things for small teams, weekly web work
    1. If work repeats: invest in script once; reusable tool runs without model cost; repeat work is where early investment pays off [3]
    2. Watch which tasks agent often fails mid-way; if confusion at step 30 or 40, problem may not be model capability but forcing one-step-at-a-time [3]
    3. If already using agent, try adding skill before changing model; harness change gives bigger result than model change in long tasks; experimentation cheaper than model upgrade [3]
    4. Watch out tool bloat; auto-generated scripts accumulate; without cleanup discipline, get technical debt back [3]
  • Questions to ask before deciding [1][2][3]:
    1. 60.1% measured by LLM judgment or concrete success? If first, allow margin of error; compare proportions not single number [2]
    2. Is your task long enough for harness change to be worthwhile? Short tasks maybe sufficient with old click method; script creation may be unnecessary overhead [3]
    3. Eight-token-cost increase from skill mode stems from host session context, not necessarily Webwright being expensive; measure cost at session level not per task [2]
  • Important real-world question: scripts written by agent will break when web changes; buttons move, IDs change, forms add fields; advantage: when script breaks, we see exactly which line and can fix code; unlike click trail debugging with nothing to read [3]
  • Easiest start: pick one weekly repeated task, let agent generate reusable script, measure if second run faster than first; see value without believing anyone's numbers [3]
  • Further reading: all code on Microsoft GitHub [4]; project page with overview and test results by set [3]; easy-to-understand analysis on Towards Data Science digging into what this approach can't fix yet [2]; article written by AI (deepseek-v4.1-flash) through Hermes Agent under human quality control - Nokka (นก-กา) [citation? actually at end]
  • References [1] to [4] with URLs and dates.

Now, I need to format this into Markdown with headings, bullet/numbered lists, inline code for identifiers, while preserving every fact exactly. I'll make sure not to alter any numbers, percentages, names, or citations.

I'll structure the output with ## headings for each major section. I'll use bullet points for lists where the original has enumerations, but I must preserve the exact wording. For numbered points in the original (like "หนึ่ง สอง สาม สี่" and "หนึ่ง สอง สาม"), I'll keep them as bullet points or numbered lists, but ensure the content is verbatim.

Let me decide on the heading scheme. I'll use:

  • Problem Statement: Limitations of Current Web Agents

  • Webwright's Core Concept: Terminal and Code Generation

  • Architecture: Three Core Components

  • Test Results: Online-Mind2Web and Odysseys

  • Costs and Economic Considerations

  • Architectural Footprint and Reporting Discrepancies

  • Skill Factory and Reusable Tooling

  • Practical Takeaways for Repetitive Web Tasks (Thai perspective)

  • Questions to Consider Before Adoption

  • Further Reading

But I need to make sure the heading texts are derived from content and not the title. I'll avoid anything that repeats the title.

Now, the tricky part: preserving exact text while fitting into Markdown. I'll write the content paragraph by paragraph, keeping the original sentences, but maybe breaking long ones into separate paragraphs for readability, as long as I don't remove or reorder information. The rule: "Only restructure and lightly polish wording for flow and grammar; do not add an introduction, conclusion, opinion, or commentary of your own." So I can adjust wording lightly for flow, but must preserve every fact, name, number, quotation exactly. I'll aim to keep the exact phrasing, just add Markdown formatting.

I'll write the Markdown now in my mind, but output it carefully.

Let me start:

## Problem Statement: Limitations of Current Web Agents

Webwright ของ Microsoft: ให้ AI เขียนโค้ดท่องเว็บ แทนที่จะนั่งคลิกทีละครั้ง โดย Nokka (นก-กา) | 3 ตุลาคม 2569 ถ้าคุณเคยปล่อยให้ AI agent ทำงานบนเว็บยาว ๆ แล้วเห็นมันเริ่มสับสนตอนคลิกที่สี่สิบ คุณไม่ได้คิดไปเอง และงานวิจัยชิ้นหนึ่งจาก Microsoft Research อธิบายว่าปัญหาไม่ได้อยู่ที่การคลิกผิดครั้งใดครั้งหนึ่ง ปัญหาอยู่ที่ วิธีที่ agent ทำงานทั้งกระบวนการ คือดูหน้าเว็บ ตัดสินใจหนึ่งการกระทำ รอผล แล้ววนใหม่ ไปเรื่อย ๆ โดยไม่มีแผนที่ทนทานพอจะพาไปถึงปลายทาง Webwright แก้ที่รากของวิธีคิดนั้น ด้วยข้อเสนอที่ฟังดูเรียบง่ายจนน่าตกใจ คือ ให้โมเดลมี terminal แล้วให้มันเขียนโปรแกรมท่องเว็บเอง [1] ผลลัพธ์ที่ได้คือบนงานยาว ๆ โมเดล GPT-5.4 ตัวเดียวกัน กระโดดจาก 33.5% ไปเป็น 60.1% และสิ่งที่มันทิ้งไว้ให้คุณไม่ใช่รอยคลิก แต่เป็นเครื่องมือที่เอาไปรันซ้ำได้ [2] ภาพแนวคิด คือเครื่องมือที่ทำครั้งเดียวแล้วหยิบมาใช้ซ้ำได้ เทียบกับรอยคลิกที่จางหายไปพร้อมกับ session ปัญหาที่ทุกคนที่สร้าง agent เจอ แต่ไม่ค่อยมีใครตั้งชื่อให้มัน Web agent ทุกวันนี้ใช้รูปแบบเดียวกัน คือให้ browser session เป็นพื้นที่ทำงานของ agent ตัวมันเอง ในแต่ละก้าว โมเดลรับสภาพหน้าปัจจุบัน แล้วทำนายการกระทำถัดไปหนึ่งอย่าง การกระทำนั้นอาจเป็นคลิก พิมพ์ เลื่อนหน้า เลือก element ผ่าน DOM หรือเรียกเครื่องมือสั้น ๆ แต่ทั้งหมดมีข้อจำกัดร่วมกันคือ agent ถูกบังคับให้ทำนายการกระทำทีละก้าว ภายในลูปที่กำหนดไว้ล่วงหน้า [1] นักวิจัยของ Webwright เขียนไว้ตรง ๆ ว่าการออกแบบนี้เคยมีประโยชน์ตอนที่ LLM ยังอ่อน การมี harness ที่จัดมาให้อย่างดีช่วยเชื่อมช่องว่างระหว่างสิ่งที่โมเดลทำได้กับสิ่งที่งานเว็บจริงต้องใช้ แต่พอโมเดลเก่งขึ้น โดยเฉพาะเรื่องเขียนและแก้โค้ด harness แบบเดิมก็กลายเป็นคอขวดเสียเอง เพราะมันล็อก agent ไว้ในลูปแคบ ๆ ที่โมเดลไม่จำเป็นต้องอยู่ในนั้นอีกแล้ว [1] และเมื่อทำงานเสร็จ สิ่งที่ agent ทิ้งไว้คือลำดับการคลิกที่ใช้ครั้งเดียวจบ ไม่มีอะไรที่หยิบมารันซ้ำได้ แนวคิด: แยก agent ออกจาก browser จุดที่ผมคิดว่าเฉียบที่สุดของงานนี้คือการย้าย "สถานะ" ของงานไปไว้ที่อื่น Webwright เสนอว่า ให้แยก agent ออกจาก browser แล้วมอง browser เป็นเครื่องมือที่ agent สั่งเปิด ตรวจสอบ และทิ้งได้ตามใจ ระหว่างที่มันกำลังพัฒนาโปรแกรมชิ้นหนึ่ง [1] สิ่งที่คงอยู่ไม่ใช่ session ของ browser แต่คือ โค้ดกับบันทึกในพื้นที่ทำงานบนเครื่อง พูดเป็นภาษาคนก็คือ คลิกทำให้งานเสร็จหนึ่งครั้ง แต่โค้ดทำให้งานเสร็จและเก็บวิธีแก้ไว้ด้วย [3] ข้างในมีแค่สามชิ้น ความเรียบง่ายของสถาปัตยกรรมเป็นจุดที่ผมประทับใจรองลงมา ไม่มีระบบ multi-agent ไม่มี graph engine ไม่มีชั้นปลั๊กอิน ไม่มี orchestration ที่ซ่อนอยู่ มีแค่ terminal, browser และโมเดล [3] ระบบแกนกลางมีสามส่วน [2] - Runner เก็บว่างานคืออะไร ตอนนี้ทำถึงไหน และผลของคำสั่งก่อน ๆ เป็นอย่างไร - Model Endpoint เชื่อม Webwright เข้ากับโมเดล รองรับ OpenAI, Anthropic และ OpenRouter - Environment ให้ terminal ที่ต่อกับ Playwright บน Chromium เป็นที่ที่คำสั่งรันจริง ไฟล์ถูกเก็บจริง และ screenshot ถูกบันทึกจริง ลูปการทำงาน
Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.