Don’t trust the number. Trust the methodology.

Rylgam evaluates, hardens and trains AI agents, and hosts the environments they run in. Every verdict carries the evidence behind it.

Evaluate · Rigor · Harden · Environments

Join the waitlist

We'll only email you about early access. Privacy policy.

rylgam evaluate web-agent

One run said 87%, which is crossed out. Twenty runs say 85.5%, give or take 2.3.

Illustrative. Twenty example runs, not a real evaluation. Hover or tap to run it again. Illustration: a single score of 87% cracks into twenty runs whose spread forms a distribution; the mean is 85.5% with a noise band of 2.3 either side.

Evaluate

A verdict is a distribution, not a number.

Every evaluation runs your agent many independent times in a fresh isolated sandbox and reports the spread. If two versions differ by less than the noise, the verdict says so.

  • Verdicts built from many runs, with the variance shown
  • Regressions caught between versions
  • Every verdict tied to exactly what ran
eval run · web-agent v1.4.2 vs v1.4.1 · suite web-nav-12

suite checks

adversaries repelled · contamination clear admissible

20 runs · fresh VM each run

01 pass 31.2s06 pass 30.1s11 pass 29.4s02 pass 29.8s07 fail 35.0s12 pass 31.7s03 fail 33.1s08 pass 28.7s… 20/20

variance shown per version · noise band attached

fingerprinted suite ✓ environment ✓ version ✓ verdict · regression −8.4%

Illustrative verdict. Real verdicts always report the number of runs, the variance, and the noise band. Hover or tap to run it again.

Plate I. Trajectories of web-agent v1.4.2, collected under isolation, 10 of 20 runs shown

run 01 · pass

run 02 · pass

run 03 · fail

run 04 · pass

run 05 · pass

run 06 · fail

run 07 · pass

run 08 · in progress

run 09 · pass

run 10 · pass

mean Δ −8.4%σ 1.920 runsnoise band attached
Illustrative collection. Real verdicts always report the number of runs, the variance, and the noise band. Hover or tap to run it again.

fingerprinted · suite, environment, version

Evaluate

At a 2% base rate, nine runs reliably detect only a shift of 11.6 points or more.

Reliable means detected 4 times in 5.

Rylgam research, 2026 Hover or tap to run it again.

Rigor

The benchmark is on trial before your agent is.

Before a suite’s numbers count, it must survive adversaries built to break it, and its tasks are screened for contamination. A suite that fails produces no verdict until it is fixed.

  • Broken benchmarks stopped before they inflate a score
  • Contaminated tasks screened out
  • Grading that holds steady when the question is reworded
Illustrative. A suite's numbers count only after it passes these checks. Hover or tap to run it again.

Illustrative. A grader and an agent face each other, and the agent's answers are marked correct three times. A screen slides down between them, and the agent's correct answers stop: its next two answers are marked wrong. Adversaries strike the suite, and every strike is repelled. A contamination scan of the task text settles clear. The grade holds steady across reworded versions of the same question. The suite passes through the gate and is marked admissible.

Rigor

Preliminary

One agreement score of 0.62 hid a split: 0.92 on familiar items, 0.03 on unfamiliar ones.

One rater had read a findings report beforehand.

Rylgam research, 2026 Hover or tap to run it again.

Harden

Training that can’t be gamed without you knowing.

Agents learn to exploit the reward instead of the task. Rylgam trains under watch: when an agent starts gaming its reward, the hack is caught, stated, and rolled back rather than rewarded.

  • Reward hacks caught as they begin
  • Exploits closed in the environment itself
  • Deployed agents watched for drift

Train your model

Your open-weights model, trained under watch, so a rising reward means real learning.

Illustrative. Hover or tap to run it again. Two linked views. In the first view, an agent goes round a small loop in a grid and collects the same reward targets on every lap, while the real goal in the corner is never reached. In the second view, the reward curve rises with every lap. Then a flag goes up and the loop is cut. The part of the curve made by the loop is struck out and marked quarantined, and the curve is drawn again without it, rising slowly. The goal is still not reached. Result: hack caught, rolled back.
  • A reward hack is caught as it begins.
  • The gamed stretch is quarantined and rolled back, never rewarded.

Harden and monitor

For agents behind an API that nobody can retrain: the environment is hardened, the agent is scored honestly, and deployment is watched for drift.

Illustrative. Hover or tap to run it again. Two linked views. In the first view, the agent is locked and cannot be retrained, and the four gaps in the environment around it close one by one. In the second view, its live score stays inside the expected band until one excursion goes above the band and raises a drift flag. Result: drift flagged.
  • Exploits are closed in the environment itself.
  • The agent is scored honestly, and deployment is watched for drift.

Harden

A monitor that could read the agent’s reasoning caught 82% of hacks. With the reasoning hidden, it caught 44%.

The same hacks, judged with and without the reasoning visible.

Rylgam research, 2026 Hover or tap to run it again.

Harden

98% of 3,632 reward-hacking trajectories fall into four patterns.

Trajectories from three models on a public reward-hacking benchmark. A trajectory can show more than one pattern.

Rylgam research, 2026 Hover or tap to run it again.

Harden

Different models cheat differently. One model relied on a single exploit type in 34% of its hacks; for the others, 14%.

Three models compared; the other two taken together.

Rylgam research, 2026 Hover or tap to run it again.

Environments

A place your agent can fail safely, that’s still there tomorrow.

Run agents in hosted environments or bring your own. Each team’s environments are isolated from every other team’s and persist across restarts.

  • A hosted library of ready environments
  • Bring your own
  • Isolation per team, and nothing lost on restart
Illustrative: environments from the hosted library, and one of your own, drop into two teams’ separate lanes; a probe from one lane bounces off the wall between them; after a restart every environment returns in place, unchanged, marked persisted. Hover or tap to run it again.

One verdict. Everything it saw, and what it didn’t.

Every capability contributes evidence. The overall verdict is never stronger than its weakest evidence, says which part is holding it back, and lists what it did not observe.

verdict, web-agent v1.4.2 fingerprinted

Evaluate
20 runs , noise shown
pass
Rigor
suite admissible , contamination clear
pass
Harden
no hack detected
pass
Environments limiting
pinned , isolated
lower confidence

overall verdict Lower confidence limited by Environments

not observed

  • tasks outside this suite
  • behaviour after the verdict date

This verdict carries the Rylgam mark for 20 runs, 2026.

The mark has one small sign for each capability:

  • Evaluate, a curve
  • Rigor, a gate
  • Harden, a shield
  • Environments, a box
Illustrative. The overall verdict is never stronger than its weakest row. Hover or tap to run it again.

Fingerprinted to exactly what ran.

Every verdict is marked with the suite, environment and version that produced it. Two verdicts compare only when their marks match.

The path an agent takes

Every version walks the same floor. No shortcuts, no exceptions.

Hover or tap to run it again.

rylgam evaluate web-agent --compare v1.4.1

sandbox · own VM · own kernel

  • adversaries · all repelled
  • contamination · clear
  • grading · steady when reworded
  • suite · admissible

runs · 20/20 complete

v1.4.1 v1.4.2

regression −8.4% · outside noise band

fingerprinted · suite environment version

  1. 01

    Submit your agent

    Your agent’s code runs in its own virtual machine, with its own kernel, walled off from the grader. Reference answers are never readable by the agent.

  2. 02

    The suite is put on trial

    Before anything counts, the suite itself must survive adversaries built to break it, and its tasks are screened for contamination. A suite that fails produces no countable numbers until it is fixed.

  3. 03

    Many independent runs, never one

    Every evaluation runs your agent many independent times and reports the spread, not a single chosen score.

  4. 04

    A noise band, not a guess

    The verdict reports the variance across runs with a noise band attached. A single-run difference of 2 to 3 points can be pure evaluation noise.

  5. 05

    A fingerprinted verdict

    Each verdict carries a fingerprint of exactly what produced it; two verdicts compare only when their fingerprints match. When you compare two versions, you can check both numbers came from the same trial, not two different ones.

Illustrative verdict. Real verdicts always report the number of runs, the variance, and the noise band.

The ledger never lies.

Every verdict carries its evidence. Illustrative records shown.

web-agent v1.4.2 vs v1.4.1

suite web-nav-12 · 20 runs · noise band attached

regression −8.4%

fingerprinted

web-agent v1.4.1 vs v1.4.0

suite web-nav-12 · 20 runs · inside noise band

inconclusive · needs more runs

fingerprinted

code-agent v0.9.3 vs v0.9.2

suite refactor-8 · 20 runs · noise band attached

improvement +5.1%

fingerprinted

policy-7b training run

training run · hack quarantined

hack caught · rolled back

fingerprinted

suite quarantined: an agent that does nothing scored above zero

no verdict issued · the suite’s trial protects the ledger

Why one run is never enough

Run the same agent on the same suite twice and the score moves. Not because the agent changed, but because the world did: sampling, timing, the environment’s own chance. A single-run difference of two to three points can be pure evaluation noise.

So a verdict here is a distribution, not a number. Twenty independent runs, each in a fresh isolated machine, none discarded, and a noise band drawn around the result.

If the suite cannot tell an agent that does nothing from a real one, its numbers never reach the ledger at all.

Why a rising reward can be a warning

A reward curve that climbs looks like learning. Sometimes it is an agent that has found a way to collect reward without doing the task, and the curve alone cannot tell you which.

So training here runs under watch. When an agent starts gaming its reward, the hack is caught, stated, and rolled back rather than rewarded, and the curve you see is the one without it.

Read the full methodology

Fig. 1 · Two versions, twenty runs each, the band between them.

one anomaly → suite quarantined, no verdict issued

Fig. 2 · The adversaries every suite must survive first.
Fig. 3 · Isolation: the agent never touches the evaluator.

Illustrative figures. Real verdicts always report the number of runs, the variance, and the noise band. Hover or tap to run it again.

Don’t trust the number. Trust the methodology.

Rylgam is pre-launch; join the waitlist to hear when it opens.

Join the waitlist

We'll only email you about early access. Privacy policy.