Skip to content

feat(evals): add Claude Opus 5.5 and GPT 6 Sol (high) results - #2432

Merged
benjamincanac merged 2 commits into
mainfrom
feat/evals-claude-opus-5.5
Sep 22, 2026
Merged

benjamincanac merged 2 commits into
mainfrom
feat/evals-claude-opus-5.5

Conversation

@benjamincanac

@benjamincanac benjamincanac commented Sep 22, 2026

Copy link
Copy Markdown
Member

Adds Claude Opus 5.5 (Claude Code) and GPT 6 Sol (high) (Codex) to public/agent-results.json with their per-eval rows. Existing experiments are untouched.

Both already pick up their icons through the claude and gpt prefixes in the evals page model map, so no code change needed.

@vercel

vercel Bot commented Sep 22, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
nuxt Ready Ready Preview Sep 22, 2026 8:16pm UTC

Request Review

@coderabbitai

coderabbitai Bot commented Sep 22, 2026

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 6fdb6dd8-ac2e-4a5b-bfcb-2563f5b9e00f

📥 Commits

Reviewing files that changed from the base of the PR and between c0fb180 and 0728ace.

📒 Files selected for processing (1)
  • public/agent-results.json

Included review availability: Your plan provides up to 2 included reviews per hour; 1 remains after this review.


📝 Walkthrough

Walkthrough

The benchmark export timestamp was updated. A new claude-opus-5.5 experiment was added with aggregate duration, cost, and pass-rate metrics. The results object now includes 31 evaluations from nuxt-000 through nuxt-ui-007. All evaluations passed on the first run except nuxt-020-fix-nuxt-module, which passed after one retry.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~10 minutes

Merge Risk: ⚪ Minimal · up to 0728a

The PR adds internally consistent benchmark results and presents no actionable merge-blocking risk.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title accurately identifies the addition of Claude Opus 5.5 results. The available changes summary does not confirm the GPT 6 Sol results, but the title remains related to the changeset.
Description check ✅ Passed The description directly relates to the changeset by describing additions to public/agent-results.json and the use of existing icon mappings.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@nuxt-com-bundle-report

nuxt-com-bundle-report Bot commented Sep 22, 2026

Copy link
Copy Markdown

Production bundle

Comparing c0fb1805 with 93d21f87. Compressed sizes are calculated from the emitted production assets.

Metric Base (Brotli) PR (Brotli) Δ Brotli Δ gzip
Client JavaScript 2.54 MiB 2.54 MiB +926 B (+0.0%) +1.2 KiB (+0.0%)
Client CSS 30.3 KiB 31.4 KiB +1.1 KiB (+3.6%) +1.4 KiB (+3.6%)
Other client assets 339.3 KiB 339.3 KiB
Total client assets 2.90 MiB 2.90 MiB +2.0 KiB (+0.1%) +2.6 KiB (+0.1%)

Largest module increases

Module Base (Brotli) PR (Brotli) Δ Brotli
/public/agent-results.json 12.3 KiB 13.2 KiB +889 B (+7.1%)
/app/pages/docs/async-data-chunk-5.js 0 B 181 B +181 B
/app/pages/blog/async-data-chunk-3.js 0 B 169 B +169 B
/app/pages/docs/async-data-chunk-4.js 0 B 168 B +168 B
/app/pages/docs/[version]/errors/async-data-chunk-8.js 0 B 153 B +153 B
/app/pages/deploy/async-data-chunk-20.js 0 B 137 B +137 B
/app/pages/blog/async-data-chunk-1.js 0 B 122 B +122 B
/app/pages/deploy/async-data-chunk-19.js 0 B 121 B +121 B
/app/pages/enterprise/agencies/async-data-chunk-7.js 0 B 120 B +120 B
/app/pages/enterprise/async-data-chunk-6.js 0 B 118 B +118 B

Module values come from Nuxt’s analyzer and are attribution estimates. This workflow is currently report-only.

Workflow run

@benjamincanac benjamincanac changed the title feat(evals): add Claude Opus 5.5 results feat(evals): add Claude Opus 5.5 and GPT 6 Sol (high) results Sep 22, 2026
@benjamincanac
benjamincanac merged commit 9dae0db into main Sep 22, 2026
17 checks passed
@benjamincanac
benjamincanac deleted the feat/evals-claude-opus-5.5 branch September 22, 2026 20:49

This branch was successfully deployed

1 active deployment
Preview 93d21f87 Deployed Sep 22, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant