web-agent v1.4.2 vs v1.4.1
regression −8.4%
fingerprinted
Rylgam evaluates, hardens and trains AI agents, and hosts the environments they run in. Every verdict carries the evidence behind it.
Evaluate · Rigor · Harden · Environments
Enter a valid email address.
Too many signups right now. Try again in a minute, or email alexhaya4@gmail.com to join.
We could not reach the waitlist. Try again in a moment, or email alexhaya4@gmail.com to join.
You are on the list.
We'll only email you about early access. Privacy policy.
rylgam evaluate web-agent
One run said 87%, which is crossed out. Twenty runs say 85.5%, give or take 2.3.
Evaluate
Every evaluation runs your agent many independent times in a fresh isolated sandbox and reports the spread. If two versions differ by less than the noise, the verdict says so.
suite checks
adversaries repelled · contamination clear admissible
20 runs · fresh VM each run
variance shown per version · noise band attached
fingerprinted suite ✓ environment ✓ version ✓ verdict · regression −8.4%
Plate I. Trajectories of web-agent v1.4.2, collected under isolation, 10 of 20 runs shown
run 01 · pass
run 02 · pass
run 03 · fail
run 04 · pass
run 05 · pass
run 06 · fail
run 07 · pass
run 08 · in progress
run 09 · pass
run 10 · pass
fingerprinted · suite, environment, version
Evaluate
At a 2% base rate, nine runs reliably detect only a shift of 11.6 points or more.
Reliable means detected 4 times in 5.
Rylgam research, 2026 Hover or tap to run it again.
Rigor
Before a suite’s numbers count, it must survive adversaries built to break it, and its tasks are screened for contamination. A suite that fails produces no verdict until it is fixed.
Illustrative. A grader and an agent face each other, and the agent's answers are marked correct three times. A screen slides down between them, and the agent's correct answers stop: its next two answers are marked wrong. Adversaries strike the suite, and every strike is repelled. A contamination scan of the task text settles clear. The grade holds steady across reworded versions of the same question. The suite passes through the gate and is marked admissible.
Rigor
Preliminary
One agreement score of 0.62 hid a split: 0.92 on familiar items, 0.03 on unfamiliar ones.
One rater had read a findings report beforehand.
Rylgam research, 2026 Hover or tap to run it again.
Harden
Agents learn to exploit the reward instead of the task. Rylgam trains under watch: when an agent starts gaming its reward, the hack is caught, stated, and rolled back rather than rewarded.
Your open-weights model, trained under watch, so a rising reward means real learning.
For agents behind an API that nobody can retrain: the environment is hardened, the agent is scored honestly, and deployment is watched for drift.
Harden
A monitor that could read the agent’s reasoning caught 82% of hacks. With the reasoning hidden, it caught 44%.
The same hacks, judged with and without the reasoning visible.
Rylgam research, 2026 Hover or tap to run it again.
Harden
98% of 3,632 reward-hacking trajectories fall into four patterns.
Trajectories from three models on a public reward-hacking benchmark. A trajectory can show more than one pattern.
Rylgam research, 2026 Hover or tap to run it again.
Harden
Different models cheat differently. One model relied on a single exploit type in 34% of its hacks; for the others, 14%.
Three models compared; the other two taken together.
Rylgam research, 2026 Hover or tap to run it again.
Environments
Run agents in hosted environments or bring your own. Each team’s environments are isolated from every other team’s and persist across restarts.
Every capability contributes evidence. The overall verdict is never stronger than its weakest evidence, says which part is holding it back, and lists what it did not observe.
verdict, web-agent v1.4.2 fingerprinted
overall verdict Lower confidence limited by Environments
not observed
This verdict carries the Rylgam mark for 20 runs, 2026.
The mark has one small sign for each capability:
Every verdict is marked with the suite, environment and version that produced it. Two verdicts compare only when their marks match.
Every version walks the same floor. No shortcuts, no exceptions.
Hover or tap to run it again.
rylgam evaluate web-agent --compare v1.4.1
sandbox · own VM · own kernel
runs · 20/20 complete
v1.4.1 v1.4.2
regression −8.4% · outside noise band
fingerprinted · suite environment version
01
Your agent’s code runs in its own virtual machine, with its own kernel, walled off from the grader. Reference answers are never readable by the agent.
02
Before anything counts, the suite itself must survive adversaries built to break it, and its tasks are screened for contamination. A suite that fails produces no countable numbers until it is fixed.
03
Every evaluation runs your agent many independent times and reports the spread, not a single chosen score.
04
The verdict reports the variance across runs with a noise band attached. A single-run difference of 2 to 3 points can be pure evaluation noise.
05
Each verdict carries a fingerprint of exactly what produced it; two verdicts compare only when their fingerprints match. When you compare two versions, you can check both numbers came from the same trial, not two different ones.
Illustrative verdict. Real verdicts always report the number of runs, the variance, and the noise band.
Every verdict carries its evidence. Illustrative records shown.
regression −8.4%
fingerprinted
inconclusive · needs more runs
fingerprinted
improvement +5.1%
fingerprinted
hack caught · rolled back
fingerprinted
Run the same agent on the same suite twice and the score moves. Not because the agent changed, but because the world did: sampling, timing, the environment’s own chance. A single-run difference of two to three points can be pure evaluation noise.
So a verdict here is a distribution, not a number. Twenty independent runs, each in a fresh isolated machine, none discarded, and a noise band drawn around the result.
If the suite cannot tell an agent that does nothing from a real one, its numbers never reach the ledger at all.
A reward curve that climbs looks like learning. Sometimes it is an agent that has found a way to collect reward without doing the task, and the curve alone cannot tell you which.
So training here runs under watch. When an agent starts gaming its reward, the hack is caught, stated, and rolled back rather than rewarded, and the curve you see is the one without it.
one anomaly → suite quarantined, no verdict issued
Illustrative figures. Real verdicts always report the number of runs, the variance, and the noise band. Hover or tap to run it again.
Rylgam is pre-launch; join the waitlist to hear when it opens.
Enter a valid email address.
Too many signups right now. Try again in a minute, or email alexhaya4@gmail.com to join.
We could not reach the waitlist. Try again in a moment, or email alexhaya4@gmail.com to join.
You are on the list.
We'll only email you about early access. Privacy policy.