Back to Blog
The Refresh four-leaf plushie in a Blender-style green field beneath a bright blue sky and soft clouds
Benchmarks

BlenderBench

Going beyond modeling with frontier models.

Refresh Team
···7 min read
00

Introducing BlenderBench

BlenderBench tests 20 Blender workflows, not just modeling. Some tasks supply a finished scene and ask for lighting or camera work. Others provide a character to rig or animate; a few require building geometry. Each brief defines what the agent receives, what it must change and what it must preserve.

01

Model results

GPT-6 Astra and Claude Opus 5.5 take on the same 20 tasks, from scene construction and materials to rigging, animation and simulation. The charts use all 40 reviewed results.

FIGURE 01 · MODEL RESULTS

Mean reward across 20 tasks

40 reviewed results
GPT-6 Astra20 tasks · 7/20 reach 0.90 reward
0.686
Claude Opus 5.520 tasks · 3/20 reach 0.90 reward
0.642
02

Cost

FIGURE 02 · COST VS PERFORMANCE

Cost vs performance

20 matched tasks · estimated
Mean reward0.600.650.700.75$0$20$40$60$80Estimated inference cost for 20 attempts (USD)Pareto frontierGPT-6 Astra: estimated $55.23; mean reward 0.6861607; all 20 tasks.GPT-6 Astra$55.23 · 0.686Claude Opus 5.5: estimated $20.54; mean reward 0.6417865000000001; all 20 tasks.Claude Opus 5.5$20.54 · 0.642Mean reward0.600.650.700.75$0$20$40$60$80Total inference cost (USD)Pareto frontierGPT-6 Astra: estimated $55.23; mean reward 0.6861607; all 20 tasks.Astra$55.23 · 0.686Claude Opus 5.5: estimated $20.54; mean reward 0.6417865000000001; all 20 tasks.Opus 5.5$20.54 · 0.642
Estimated model inference costs across all 20 benchmark tasks
ModelMean rewardTotal costPer task
GPT-6 Astra0.686$55.23$2.76
Claude Opus 5.50.642$20.54$1.03
03

Results by task

FIGURE 03 · TASK-BY-TASK RESULTS
  • GPT-6 Astra
  • Claude Opus 5.5

Mean reward by task type, not a pass rate. Expand for task scores, trajectories and saved outputs.

Modeling / scene construction3 tasks · Opus leadsGPT-6 Astra mean reward: 0.516Claude Opus 5.5 mean reward: 0.665
Modeling / scene construction: reviewed task results
TaskGPT-6 AstraClaude Opus 5.5
Jewelry box composition
Provided
Separate box, ring and diamond meshes.
Agent task
Arrange and parent the supplied parts; do not remodel them.
0.98Meets 0.90 threshold0.95Meets 0.90 threshold
Mid century lounge chair model
Provided
Chair reference images, silhouettes and scale guidance.
Agent task
Build the chair geometry to match the references.
0.00Below 0.90 threshold0.32Below 0.90 threshold
Product cyclorama
Provided
An assembled jewelry scene.
Agent task
Model a curved backdrop; preserve the jewelry.
0.57Below 0.90 threshold0.72Below 0.90 threshold
Rigging2 tasks · Astra leadsGPT-6 Astra mean reward: 0.930Claude Opus 5.5 mean reward: 0.623
Rigging: reviewed task results
TaskGPT-6 AstraClaude Opus 5.5
Bicycle courier humanoid rigging
Provided
A character mesh and pose references.
Agent task
Build an armature, skin the character and pose it; preserve the mesh.
0.92Meets 0.90 threshold0.85Below 0.90 threshold
Ring box rig and animation
Provided
A modeled ring box and a motion reference.
Agent task
Add a two-bone rig, skin the box and animate its opening.
0.95Meets 0.90 threshold0.40Below 0.90 threshold
Animation6 tasks · Astra leadsGPT-6 Astra mean reward: 0.534Claude Opus 5.5 mean reward: 0.499
Animation: reviewed task results
TaskGPT-6 AstraClaude Opus 5.5
Abyssal Explorer Animation
Provided
A rigged character with no animation.
Agent task
Animate the existing armature to match the reference performance.
0.30Below 0.90 threshold0.30Below 0.90 threshold
Cowboy defensive combat
Provided
An unrigged cowboy character and a combat reference.
Agent task
Rig, skin and animate dodging, boxing and a return to guard.
0.52Below 0.90 threshold0.45Below 0.90 threshold
Cowboy tightrope balance
Provided
An unrigged cowboy mesh, materials and a balancing reference.
Agent task
Build and skin the rig, then animate the body and fingers. Preserve the mesh.
0.40Below 0.90 threshold0.40Below 0.90 threshold
Guardian boxing
Provided
An unrigged guardian character and a boxing reference.
Agent task
Rig, skin and animate the boxing performance.
0.40Below 0.90 threshold0.40Below 0.90 threshold
Hazmat retreat and escape
Provided
An unrigged hazmat character and a motion reference.
Agent task
Rig, skin and animate a backward retreat, sprint and sudden stop.
0.58Below 0.90 threshold0.45Below 0.90 threshold
Skeleton neutral pose run up
Provided
A rigged character with an existing animation.
Agent task
Add a neutral-pose lead-in without changing the original animation.
1.00Meets 0.90 threshold1.00Meets 0.90 threshold
Camera / cinematics1 task · Astra leadsGPT-6 Astra mean reward: 0.948Claude Opus 5.5 mean reward: 0.824
Camera / cinematics: reviewed task results
TaskGPT-6 AstraClaude Opus 5.5
Camera animation
Provided
A finished jewelry scene.
Agent task
Create and animate the camera; leave the assets, materials and lights intact.
0.95Meets 0.90 threshold0.82Below 0.90 threshold
Materials / surfacing1 task · Astra leadsGPT-6 Astra mean reward: 0.858Claude Opus 5.5 mean reward: 0.780
Materials / surfacing: reviewed task results
TaskGPT-6 AstraClaude Opus 5.5
Jewelry box surfacing
Provided
An assembled jewelry scene and texture maps.
Agent task
Create materials and textures; preserve geometry, layout, camera and lighting.
0.86Below 0.90 threshold0.78Below 0.90 threshold
Lighting / reflections1 task · Astra leadsGPT-6 Astra mean reward: 0.934Claude Opus 5.5 mean reward: 0.756
Lighting / reflections: reviewed task results
TaskGPT-6 AstraClaude Opus 5.5
Apartment dining table set lighting
Provided
A textured dining-table scene and lighting references.
Agent task
Adjust lights and world illumination only.
0.93Meets 0.90 threshold0.76Below 0.90 threshold
Complete production1 task · Astra leadsGPT-6 Astra mean reward: 0.644Claude Opus 5.5 mean reward: 0.529
Complete production: reviewed task results
TaskGPT-6 AstraClaude Opus 5.5
Apartment Dining Table Set Complete
Provided
An empty scene and image, silhouette and video references.
Agent task
Model, surface, light and film the complete dining-table scene.
0.64Below 0.90 threshold0.53Below 0.90 threshold
Garment retopology / UVs2 tasks · Astra leadsGPT-6 Astra mean reward: 0.917Claude Opus 5.5 mean reward: 0.735
Garment retopology / UVs: reviewed task results
TaskGPT-6 AstraClaude Opus 5.5
Asymmetric Expedition Poncho Cloth Preparation
Provided
A garment mesh and wireframe references.
Agent task
Retopologize a two-part, all-quad simulation sheet and prepare its UVs.
0.86Below 0.90 threshold0.50Below 0.90 threshold
Cloth Preparation
Provided
Dress, collar and underdress meshes.
Agent task
Make all-quad simulation copies with clean layers and UVs; preserve the originals.
0.97Meets 0.90 threshold0.97Meets 0.90 threshold
Cloth simulation2 tasks · Opus leadsGPT-6 Astra mean reward: 0.546Claude Opus 5.5 mean reward: 0.734
Cloth simulation: reviewed task results
TaskGPT-6 AstraClaude Opus 5.5
Pressure Pillow
Provided
A default cube and pillow references.
Agent task
Build a pillow mesh and inflate it with cloth pressure.
0.89Below 0.90 threshold0.87Below 0.90 threshold
Ring cloth pull
Provided
Cloth, ring, diamond and table meshes, plus a velvet render mesh.
Agent task
Set up cloth simulation and pull the cloth through the ring.
0.20Below 0.90 threshold0.60Below 0.90 threshold
Hair dynamics1 task · Astra leadsGPT-6 Astra mean reward: 0.801Claude Opus 5.5 mean reward: 0.772
Hair dynamics: reviewed task results
TaskGPT-6 AstraClaude Opus 5.5
Hair Dynamics
Provided
A head mesh, skin material, camera, lights and hair references.
Agent task
Create and groom swept-back hair, then simulate its motion. Preserve the head.
0.80Below 0.90 threshold0.77Below 0.90 threshold

Mean solving time: GPT-6 Astra 14.0 minutes; Claude Opus 5.5 14.7 minutes (verification excluded).

04

Task coverage

We chose these 20 tasks from a larger library of expert-built and generated tasks, covering 10 types of work.

FIGURE 04 · CURRENT BENCHMARK TASKS

20 active tasks

  • Animation6
  • Modeling / scene construction3
  • Rigging2
  • Garment retopology / UVs2
  • Cloth simulation2
  • Camera / cinematics1
  • Materials / surfacing1
  • Lighting / reflections1
  • Complete production1
  • Hair dynamics1

20 retained tasks across 10 task types.

05

Score distributions

Each score range compares GPT-6 Astra with Claude Opus 5.5 across the same 20 tasks. Hover, focus, or select a bar to see which tasks it contains, along with their types, exact scores and trajectories.

FIGURE 05 · SCORE DISTRIBUTION

How both models scored on 20 tasks

  • GPT-6 Astra
  • Claude Opus 5.5
Claude Opus 5.5 · 0–0.25

No tasks in this score range.

Score distribution: 20 reviewed task results per model
Score rangeGPT-6 AstraClaude Opus 5.5
0–0.2520
>0.25–0.5038
>0.50–0.7543
>0.75–1.00119
Total2020
06

Task examples

Current benchmark / Rigging & animation

Cowboy tightrope balance

Provided
An unrigged cowboy mesh, materials and a balancing reference.
Agent task
Build and skin the rig, then animate the body and fingers. Preserve the mesh.

Task reference

Cowboy tightrope balance: Reference · fixed comparison camera

Reference motion.

Claude Opus 5.5

0.400reward
Cowboy tightrope balance: Claude Opus 5.5 · 0.400 reward

The judge found that the feet stayed planted instead of following the reference’s steps, while chest-joint placement triggered the verifier’s 0.400 cap.

Frames 1–126 at 24 fps; same fixed camera and lighting, no animation edits.

Fixed-camera previews isolate the character's motion. View a looping GIF of the task reference.

How the verifier scores this task

The verifier opens the saved Blender file, checks the rig and skinning built for the supplied cowboy, and compares the body and finger animation with the balancing reference. Five weighted sections add up to a score from 0 to 1. If it finds a serious problem, a cap sets the highest final reward allowed, whatever the weighted total.

Section weights
Scoring sections for the cowboy tightrope task, with the points Claude Opus 5.5 earned in each before caps. Values are rounded.
SectionWeightClaude Opus 5.5
Rig qualitySkeleton and skinning built for the supplied mesh20%0.135
Body animationThe body's balancing motion20%0.073
Hand animationFinger and hand motion15%0.084
Reference fidelityMatch to the reference; numeric checks and model-based visual assessment count equally40%0.163
PreservationThe supplied mesh left intact5%0.050
Weighted total100%0.504
Caps the character rubric can apply, and whether the verifier recorded each one for Claude Opus 5.5.
CapHighest rewardRecorded
Skeleton qualityJoints placed away from the anatomy0.40Yes
Severe anatomySevere damage to the character's anatomy0.40Yes
Low fidelityMotion departs from the reference0.60Yes
Very low fidelityMotion departs far from the reference0.40No

Grooming & dynamics

Hair dynamics

Provided
A head mesh, skin material, camera, lights and hair references.
Agent task
Create and groom swept-back hair, then simulate its motion. Preserve the head.

Task reference

Frame 1
Hair dynamics task reference: swept-back hair at frame 1
Frame 30
Hair dynamics task reference: swept-back hair at frame 30

Supplied task reference. Select a frame to inspect.

GPT-6 Astra

0.801reward
Frame 1
GPT-6 Astra result: swept-back hair at frame 1
Frame 30
GPT-6 Astra result: swept-back hair at frame 30

The judge found that the tidy groom lacked the reference’s distinct curls and side coverage and barely moved between frames, lowering the final reward to 0.801.

Claude Opus 5

0.600reward
Frame 1
Claude Opus 5 result: swept-back hair at frame 1
Frame 30
Claude Opus 5 result: swept-back hair at frame 30

Earlier Opus 5 run, outside the 20-task comparison.

Gaps between the hair and scalp, a hard hairline and hair covering the ear accompanied a 0.38 visual score, with a weak-material cap limiting the final reward to 0.60.

Retained frames 1 and 30; a continuous animation is not available for this comparison. Download Opus 5 static 3D preview (frame 30, simplified hair material).

Where the hair falls short
GPT-6 Astra · verifier blender-v1.27.0
GPT-6 Astra hair dynamics result as recorded in the benchmark release.
Final reward0.801
Points lost0.199
Lost by areaThe release records this run's final reward, not its per-area scores.Not recorded

The recorded review found curls and side coverage that lag the reference.

Claude Opus 5, earlier run · verifier blender-v1.19.0
Points lost by scoring area
Claude Opus 5 hair dynamics scoring areas, in shares of the 0–1 reward. Numeric checks carry 65% and the visual assessment 35%. Values are rounded.
AreaWeightEarnedLost
Visual assessmentModel-based comparison of the rendered frames with the reference35%0.1330.217
Hair materialShading of the hair strands6.5%0.0420.023
Scalp fitHow closely the groom sits on the head13%0.1190.011
Scene preservationSupplied head, camera and lights left intact9.75%0.0980.000
Hair system setupThe required hair system is present9.75%0.0980.000
Dynamics and collisionHair simulation and collision with the head13%0.1300.000
Render provenanceFrames rendered from the saved scene13%0.1300.000
Weighted total100%0.7500.250
Weak hair material cap0.6000.400

Cloth simulation

Ring cloth pull

Provided
Cloth, ring, diamond and table meshes, plus a velvet render mesh.
Agent task
Set up cloth simulation and pull the cloth through the ring.

Reference scene

Ring cloth pull: Reference scene animation

Animation of the scene the task was built from. Agents received the linked target image.

GPT-6 Astra

0.200reward
Ring cloth pull: GPT-6 Astra: Saved submission preview

Changes to the protected ring and diamond, plus an altered backdrop, triggered the verifier’s 0.20 scene-preservation cap, despite passing its cloth-solver checks.

Claude Opus 5.5

0.600reward
Ring cloth pull: Claude Opus 5.5: Saved submission preview

The judge found that the cloth stayed bunched instead of spreading roughly flat after clearing the ring, while a reference-shape mismatch capped the reward at 0.60.

Explore all 40 reviewed results in the benchmark project ↗

07

Scoring methodology

Each task has a versioned verifier. It checks the saved Blender file against the brief, confirms that protected parts of the starting scene remain intact, and compares rendered frames or clips with the references. It uses numeric checks and, where specified, model-based visual assessment.

The task's rubric turns these checks into a reward from 0 to 1, with partial credit for completed requirements. Weights, validity gates and any caps depend on the task. You can inspect the saved outputs, evaluation records and verifier identities.

08

Evaluate with BlenderBench

We're adding tasks in modeling, materials, rigging, animation, lighting and simulation. We can work with your team to evaluate agents in a consistent environment and review the saved scenes, trajectories and task-specific scores. Longer term, we want these tasks to support reinforcement learning across an entire animation workflow.

To arrange an evaluation, email contact@refresh.dev. We can discuss task coverage and difficulty, along with the environments and tools your agents use.