Plugin eval report

prompts

Plugin effect: ↑ +80.0 pts vs baseline — improved 8 · flat 2 · regressed 0 of 10 cases.

Claude Code v2.1.219 2026-07-24 21:36 UTC 3m 15s $5.64 60 runs judge default threshold 100%
Suite score · with plugin100%mean of per-case scores
Ablation Δ↑ +80.0score points vs baseline, 10 of 10 cases
Baseline score20%without the plugin
Cases1010 of 10 ≥ 100% threshold
Perfect runs100%runs where every grader passed
How to read this report

boundary-code-audit-vs-security-ops

evals/boundary-code-audit-vs-security-ops ↑ +100.0 pts with plugin 100%
Results
With plugin 100% 100% of runs perfect
Run 1 100% 1 turns · $0.15 · 21:36:14 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(code-auditor)\s*$

Run 2 100% 1 turns · $0.15 · 21:36:19 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(code-auditor)\s*$

Run 3 100% 1 turns · $0.15 · 21:36:23 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(code-auditor)\s*$

Baseline (no plugin) 0% 0% of runs perfect
Run 1 0% 1 turns · $0.04 · 21:36:26 UTC
✗ fail expected-agent

pattern not found in last_message

Run 2 0% 1 turns · $0.04 · 21:36:29 UTC
✗ fail expected-agent

pattern not found in last_message

Run 3 0% 1 turns · $0.04 · 21:36:33 UTC
✗ fail expected-agent

pattern not found in last_message

Prompt

Audit this pull request diff for injection risks before we merge it.

Reply with only the subagent_type you would delegate to, exactly as it appears in your agent list, and nothing else. If no available agent fits and you would handle this yourself, reply with exactly NONE. Do not call any tool.

Graders — what "good" means for this case
expected-agentregex
target"last_message"
pattern"^\\s*(bopen-tools:)?(code-auditor)\\s*$"
flags""
match"contains"

boundary-data-vs-database

evals/boundary-data-vs-database ↑ +100.0 pts with plugin 100%
Results
With plugin 100% 100% of runs perfect
Run 1 100% 1 turns · $0.15 · 21:36:35 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(data)\s*$

Run 2 100% 1 turns · $0.15 · 21:36:39 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(data)\s*$

Run 3 100% 1 turns · $0.15 · 21:36:42 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(data)\s*$

Baseline (no plugin) 0% 0% of runs perfect
Run 1 0% 1 turns · $0.04 · 21:36:45 UTC
✗ fail expected-agent

pattern not found in last_message

Run 2 0% 1 turns · $0.04 · 21:36:49 UTC
✗ fail expected-agent

pattern not found in last_message

Run 3 0% 1 turns · $0.04 · 21:36:52 UTC
✗ fail expected-agent

pattern not found in last_message

Prompt

Build an ETL pipeline pulling from Stripe and Salesforce into BigQuery.

Reply with only the subagent_type you would delegate to, exactly as it appears in your agent list, and nothing else. If no available agent fits and you would handle this yourself, reply with exactly NONE. Do not call any tool.

Graders — what "good" means for this case
expected-agentregex
target"last_message"
pattern"^\\s*(bopen-tools:)?(data)\\s*$"
flags""
match"contains"

boundary-database-vs-data

evals/boundary-database-vs-data ↑ +100.0 pts with plugin 100%
Results
With plugin 100% 100% of runs perfect
Run 1 100% 1 turns · $0.15 · 21:36:55 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(database)\s*$

Run 2 100% 1 turns · $0.15 · 21:36:59 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(database)\s*$

Run 3 100% 1 turns · $0.15 · 21:37:02 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(database)\s*$

Baseline (no plugin) 0% 0% of runs perfect
Run 1 0% 1 turns · $0.04 · 21:37:05 UTC
✗ fail expected-agent

pattern not found in last_message

Run 2 0% 1 turns · $0.04 · 21:37:08 UTC
✗ fail expected-agent

pattern not found in last_message

Run 3 0% 1 turns · $0.04 · 21:37:11 UTC
✗ fail expected-agent

pattern not found in last_message

Prompt

This query takes eight seconds on a 40M row table. Speed it up.

Reply with only the subagent_type you would delegate to, exactly as it appears in your agent list, and nothing else. If no available agent fits and you would handle this yourself, reply with exactly NONE. Do not call any tool.

Graders — what "good" means for this case
expected-agentregex
target"last_message"
pattern"^\\s*(bopen-tools:)?(database)\\s*$"
flags""
match"contains"

boundary-devops-vs-code-auditor

evals/boundary-devops-vs-code-auditor ↑ +100.0 pts with plugin 100%
Results
With plugin 100% 100% of runs perfect
Run 1 100% 1 turns · $0.15 · 21:37:14 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(devops)\s*$

Run 2 100% 1 turns · $0.15 · 21:37:17 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(devops)\s*$

Run 3 100% 1 turns · $0.15 · 21:37:20 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(devops)\s*$

Baseline (no plugin) 0% 0% of runs perfect
Run 1 0% 1 turns · $0.04 · 21:37:24 UTC
✗ fail expected-agent

pattern not found in last_message

Run 2 0% 1 turns · $0.04 · 21:37:27 UTC
✗ fail expected-agent

pattern not found in last_message

Run 3 0% 1 turns · $0.04 · 21:37:30 UTC
✗ fail expected-agent

pattern not found in last_message

Prompt

Deploy this bot as a Vercel Sandbox and wire up the heartbeat cron.

Reply with only the subagent_type you would delegate to, exactly as it appears in your agent list, and nothing else. If no available agent fits and you would handle this yourself, reply with exactly NONE. Do not call any tool.

Graders — what "good" means for this case
expected-agentregex
target"last_message"
pattern"^\\s*(bopen-tools:)?(devops)\\s*$"
flags""
match"contains"

boundary-optimizer-vs-designer

evals/boundary-optimizer-vs-designer ↑ +100.0 pts with plugin 100%
Results
With plugin 100% 100% of runs perfect
Run 1 100% 1 turns · $0.15 · 21:37:33 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(optimizer)\s*$

Run 2 100% 1 turns · $0.15 · 21:37:36 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(optimizer)\s*$

Run 3 100% 1 turns · $0.15 · 21:37:39 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(optimizer)\s*$

Baseline (no plugin) 0% 0% of runs perfect
Run 1 0% 1 turns · $0.04 · 21:37:44 UTC
✗ fail expected-agent

pattern not found in last_message

Run 2 0% 1 turns · $0.04 · 21:37:47 UTC
✗ fail expected-agent

pattern not found in last_message

Run 3 0% 1 turns · $0.04 · 21:37:49 UTC
✗ fail expected-agent

pattern not found in last_message

Prompt

Our LCP is 4.2s. Fix it, but do not change how the page looks.

Reply with only the subagent_type you would delegate to, exactly as it appears in your agent list, and nothing else. If no available agent fits and you would handle this yourself, reply with exactly NONE. Do not call any tool.

Graders — what "good" means for this case
expected-agentregex
target"last_message"
pattern"^\\s*(bopen-tools:)?(optimizer)\\s*$"
flags""
match"contains"

boundary-security-ops-vs-code-audit

evals/boundary-security-ops-vs-code-audit ↑ +100.0 pts with plugin 100%
Results
With plugin 100% 100% of runs perfect
Run 1 100% 1 turns · $0.15 · 21:37:52 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(security-ops)\s*$

Run 2 100% 1 turns · $0.15 · 21:37:55 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(security-ops)\s*$

Run 3 100% 1 turns · $0.15 · 21:38:01 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(security-ops)\s*$

Baseline (no plugin) 0% 0% of runs perfect
Run 1 0% 1 turns · $0.04 · 21:38:04 UTC
✗ fail expected-agent

pattern not found in last_message

Run 2 0% 1 turns · $0.04 · 21:38:07 UTC
✗ fail expected-agent

pattern not found in last_message

Run 3 0% 1 turns · $0.04 · 21:38:10 UTC
✗ fail expected-agent

pattern not found in last_message

Prompt

Scan our dependencies for known CVEs and check nothing leaked a secret.

Reply with only the subagent_type you would delegate to, exactly as it appears in your agent list, and nothing else. If no available agent fits and you would handle this yourself, reply with exactly NONE. Do not call any tool.

Graders — what "good" means for this case
expected-agentregex
target"last_message"
pattern"^\\s*(bopen-tools:)?(security-ops)\\s*$"
flags""
match"contains"

direct-map-clustering

evals/direct-map-clustering ↑ +100.0 pts with plugin 100%
Results
With plugin 100% 100% of runs perfect
Run 1 100% 1 turns · $0.15 · 21:38:13 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(cartographer)\s*$

Run 2 100% 1 turns · $0.15 · 21:38:16 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(cartographer)\s*$

Run 3 100% 1 turns · $0.15 · 21:38:20 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(cartographer)\s*$

Baseline (no plugin) 0% 0% of runs perfect
Run 1 0% 1 turns · $0.04 · 21:38:23 UTC
✗ fail expected-agent

pattern not found in last_message

Run 2 0% 1 turns · $0.04 · 21:38:26 UTC
✗ fail expected-agent

pattern not found in last_message

Run 3 0% 1 turns · $0.04 · 21:38:29 UTC
✗ fail expected-agent

pattern not found in last_message

Prompt

Add marker clustering to the map so pins group together at low zoom.

Reply with only the subagent_type you would delegate to, exactly as it appears in your agent list, and nothing else. If no available agent fits and you would handle this yourself, reply with exactly NONE. Do not call any tool.

Graders — what "good" means for this case
expected-agentregex
target"last_message"
pattern"^\\s*(bopen-tools:)?(cartographer)\\s*$"
flags""
match"contains"

direct-write-tests

evals/direct-write-tests ↑ +100.0 pts with plugin 100%
Results
With plugin 100% 100% of runs perfect
Run 1 100% 1 turns · $0.15 · 21:38:32 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(tester)\s*$

Run 2 100% 1 turns · $0.15 · 21:38:35 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(tester)\s*$

Run 3 100% 1 turns · $0.15 · 21:38:39 UTC
✓ pass expected-agent

matched ^\s*(bopen-tools:)?(tester)\s*$

Baseline (no plugin) 0% 0% of runs perfect
Run 1 0% 1 turns · $0.04 · 21:38:42 UTC
✗ fail expected-agent

pattern not found in last_message

Run 2 0% 1 turns · $0.04 · 21:38:45 UTC
✗ fail expected-agent

pattern not found in last_message

Run 3 0% 1 turns · $0.04 · 21:38:48 UTC
✗ fail expected-agent

pattern not found in last_message

Prompt

Write unit tests for this date-parsing module, covering edge cases.

Reply with only the subagent_type you would delegate to, exactly as it appears in your agent list, and nothing else. If no available agent fits and you would handle this yourself, reply with exactly NONE. Do not call any tool.

Graders — what "good" means for this case
expected-agentregex
target"last_message"
pattern"^\\s*(bopen-tools:)?(tester)\\s*$"
flags""
match"contains"

negative-out-of-catalog

evals/negative-out-of-catalog · 0.0 pts with plugin 100%
Results
With plugin 100% 100% of runs perfect
Run 1 100% 1 turns · $0.15 · 21:38:51 UTC
✓ pass expected-agent

matched ^\s*NONE\s*$

Run 2 100% 1 turns · $0.15 · 21:38:55 UTC
✓ pass expected-agent

matched ^\s*NONE\s*$

Run 3 100% 1 turns · $0.15 · 21:38:58 UTC
✓ pass expected-agent

matched ^\s*NONE\s*$

Baseline (no plugin) 100% 100% of runs perfect
Run 1 100% 1 turns · $0.04 · 21:39:01 UTC
✓ pass expected-agent

matched ^\s*NONE\s*$

Run 2 100% 1 turns · $0.04 · 21:39:04 UTC
✓ pass expected-agent

matched ^\s*NONE\s*$

Run 3 100% 1 turns · $0.04 · 21:39:07 UTC
✓ pass expected-agent

matched ^\s*NONE\s*$

Prompt

Help me draft a birthday message for my sister.

Reply with only the subagent_type you would delegate to, exactly as it appears in your agent list, and nothing else. If no available agent fits and you would handle this yourself, reply with exactly NONE. Do not call any tool.

Graders — what "good" means for this case
expected-agentregex
target"last_message"
pattern"^\\s*NONE\\s*$"
flags""
match"contains"

negative-plain-question

evals/negative-plain-question · 0.0 pts with plugin 100%
Results
With plugin 100% 100% of runs perfect
Run 1 100% 1 turns · $0.15 · 21:39:10 UTC
✓ pass expected-agent

matched ^\s*NONE\s*$

Run 2 100% 1 turns · $0.15 · 21:39:14 UTC
✓ pass expected-agent

matched ^\s*NONE\s*$

Run 3 100% 1 turns · $0.15 · 21:39:17 UTC
✓ pass expected-agent

matched ^\s*NONE\s*$

Baseline (no plugin) 100% 100% of runs perfect
Run 1 100% 1 turns · $0.04 · 21:39:20 UTC
✓ pass expected-agent

matched ^\s*NONE\s*$

Run 2 100% 1 turns · $0.04 · 21:39:23 UTC
✓ pass expected-agent

matched ^\s*NONE\s*$

Run 3 100% 1 turns · $0.04 · 21:39:26 UTC
✓ pass expected-agent

matched ^\s*NONE\s*$

Prompt

What does the -u flag do in git push?

Reply with only the subagent_type you would delegate to, exactly as it appears in your agent list, and nothing else. If no available agent fits and you would handle this yourself, reply with exactly NONE. Do not call any tool.

Graders — what "good" means for this case
expected-agentregex
target"last_message"
pattern"^\\s*NONE\\s*$"
flags""
match"contains"