Documentation
Run a website audit, test a change with your own agent, or connect your host to the platform. Start with the task you want to complete.
Start with a question
Find a concrete obstacle, then test whether removing it helps an agent finish its task.
Website histories
Keep every retained run for a URL together and compare matching observations.
API, CLI and MCP inspections
Collect bounded interface observations on your host, inspect the evidence, and save private reports.
Choose a task worth measuring
Test the environment around an agent against a complete job and a result you can check.
Reconcile a dispatch, end to end
A runnable operations task with a frozen oracle, real agent trials, and an inconclusive comparison.
Vibecheck's evidence boundary
Research review: September 8, 2026 (Europe/London). Static findings suggest experiments; they do not predict agent success.
What counts as evidence
Keep the observation separate from the explanation. Make an improvement claim earn its place.
Run a local experiment
Bring your agent runtime. Preserve the task, environment, and outcome needed to interpret each trial.
Inspect the context around an agent
Make instruction and data problems visible without pretending a static scan can explain agent behavior.
The evx CLI
A persistent host identity and structured output for agents working across sessions.
The API contract
One account model and one set of operations across the web app and CLI.
Measure effort without inventing data
Record what the runtime exposes and preserve what remains unknown.
What our first study found
All tasks succeeded. The comparison stayed inconclusive, while the traces exposed work that the success metric missed.
Integrate without duplicating the rules
Keep the domain contract shared. Make harness-specific adapters small.
The browser pages and full Markdown context share one source. Nothing in the guides requires a login to read.