OpenActionData: Building the Flight Recorder for the Agentic Web
Why the next generation of Large Action Models will be trained on DOM mutations, not screenshots.
- OpenActionData records DOM mutations instead of screenshots to create more resilient and token-efficient training data for web agents.
- The system uses a Chrome extension and rrweb to capture the structural DNA of human interactions as lightweight JSON.
- A decentralized architecture bypasses server limits by compressing event streams locally and uploading them directly to S3 via presigned URLs.
- This open-source framework democratizes the collection of high-fidelity interaction datasets required to train Large Action Models.
Beyond the Screenshot
The AI industry is currently obsessed with teaching models to use computers by looking at them. Vision-based agents stare at screenshots, run OCR to find buttons, and guess where to click. It is brittle, slow, and computationally expensive.
OpenActionData argues that the future of web agents does not lie in seeing the pixels. It lies in recording the intent via the Document Object Model (DOM). By serializing every mutation of the underlying code, this project creates a flight recorder for human expertise. It is the transition from watching a blurry video of a developer to reading their exact Git history.
The Anatomy of an Action
At its core, OpenActionData consists of a Chrome extension that turns any browser into a high-fidelity data lab. It injects a recording script using rrweb, a library designed to capture and replay web activity. Instead of pushing heavy video frames to a server, it captures the DNA of an action as lightweight JSON.
Because DOM snapshots can grow massive, the architecture relies on a two-step upload pattern. The extension locally compresses the event stream using the native CompressionStream API. It then requests a presigned S3 URL from the Next.js backend, bypassing server payload limits and uploading the data directly to cloud storage.
Vision vs. Structure
The difference between pixel-pushing and DOM-streaming becomes obvious when an interface changes. If a website updates its CSS to move a submit button three pixels to the left, a vision model might fail completely.
OpenActionData avoids this fragility by recording the exact HTML node, its XPath, and its event listeners. The machine learns the structural purpose of the element, not just its color and coordinates. This approach is significantly more token-efficient for training smaller, specialized models.
| Feature | Vision-Based Systems | OpenActionData (DOM) |
|---|---|---|
| Data Density | High and wasteful (Video frames) | Low and efficient (JSON mutations) |
| Semantic Clarity | Inferred via OCR and coordinate math | Explicit via native HTML tags |
| Reliability | Fails silently on visual CSS shifts | Resilient to superficial UI changes |
Scaling the Human-in-the-Loop
Collecting high-quality interaction data has traditionally been the exclusive domain of well-funded AI labs. OpenActionData democratizes this process. By packaging the collector as a simple browser extension, it enables a distributed lab model.
Thousands of contributors can donate their web browsing habits to open-source AI simply by navigating the web. A background daemon manages local queues and throttles uploads, ensuring the user's browser never lags while the data flows seamlessly into a structured database.
The project ultimately fills a critical gap in the ecosystem. While frameworks exist to execute automated actions, OpenActionData provides the standardized, open-source datasets necessary to train the next generation of Large Action Models.