Vibecheck's evidence boundary
Research review: September 8, 2026 (Europe/London). Static findings suggest experiments; they do not predict agent success.
Read as MarkdownMeasure the interface the agent actually receives
Modern agents can receive a browser accessibility tree, screenshots, transformed page content, or structured tool responses. Raw HTML length is a different measurement. Vibecheck inspects a fetched document and origin companion files; it does not observe those other representations or complete a task.
The BrowserGym implementation provides task-based browser evaluation infrastructure. It supports the need for explicit observation and outcome contracts; it does not validate Vibecheck's historical HTML ratios or numerical cutoffs.
Optional conventions need a named consumer
Google's current generative search guidance says llms.txt and special structured data are not prerequisites for its generative search features. That is evidence about Google Search, not every agent. Before adding such a file to improve an agent workflow, establish whether the intended client requests and uses it.
Chrome's WebMCP announcement describes an early preview for explicit browser tools. This is a promising experiment surface, not evidence that every agent supports it or that the advertised tools are correct and safe.
Context design remains an experimental question
The June 2026 revision of Evaluating AGENTS.md challenges the assumption that adding repository context generally improves coding outcomes. A July two-agent ablation study also found no measurable correctness benefit within its reported task set and bounds. These preprints do not establish that useful instructions should always be removed.
Anthropic's advanced tool-use engineering report describes benefits and tradeoffs from deferred discovery and programmatic orchestration on its workloads. Discovery adds work too. Test your actual tool library, preserve access to full context, and check for instructions that contradict each other.
Read findings with their limits
Accessibility standards provide a basis for inspecting names and semantics, but our bounded static checks do not certify WCAG conformance or reproduce a browser's complete accessible-name algorithm. The current Accessible Name 1.2 text is a working draft.
Historical scores retain uncalibrated checklist weights and thresholds. A higher score is not evidence that an agent became more effective. Keep the report version with the observation, inspect missing coverage, and use independently verified comparisons for claims about task outcomes.