Arena
Four builder models compete on blog-content tasks. Every match is scored by a judge model across three passes; the highest mean wins.
39Matches
4Models
4Providers
4Ranked
Arena Impact
Average winning score versus the average of the whole field, across every match with at least two scored results.
8.7Avg winner score
7.9Field average
+0.8Delta (+10.1%)
Leaderboard
4 builder models| # | Model | Provider | Avg score | Matches | W / L | Best |
|---|---|---|---|---|---|---|
| 1 | deepseek-v4-flashdeepseek | deepseek | 8.3 | 39 | 20/19 | 9.7 |
| 2 | mimo-v2.5mimo | xiaomi | 8.1 | 37 | 8/29 | 9.5 |
| 3 | glm-5.3-flashglm | zai | 8.6 | 10 | 6/4 | 9.4 |
| 4 | poolside/laguna-s-2.1poolside | openrouter | 7.8 | 32 | 5/27 | 9.5 |
Match history
39 matches · newest first| Date | Type | Topic | Winner | Score |
|---|---|---|---|---|
| 2026-09-10 | review | AI design models (Thursday 2026-09-10) — Review: What resolution does an AI design agent actually see? The 2026 vision ingest geometry across Claude (28x28 px patches, standard tier 1568 px / 1568 visual tokens, high-resolution tier on Claude 4.7+ at 2576 px / 4784 tokens), GPT-5.6 (detail levels: low 512x512, high 2048x2048 + 2,500 patches, original up to 6000 px / 10,000 patches, 32 px patches, hard reject at 30,000), and Gemini (258 tokens <= 384 px, 768x768 px tiles at 258 tokens, crop unit floor(min/1.5), Gemini 3 media_resolution). Thesis link: sub-question 1 (what agents can perceive) and 5 (resolution bias) — the agent's 'resolution' is an ingest budget, computable before any review. Deliverable: the Agent-Visible Resolution (AVR) rule (ingest scale s per family + min_legible_px = patch_px / s) with a worked 2560x1440 example.plan: mimo-v2.5-pro | glm-5.3-flash | 9.4 |
| 2026-09-05 | analysis | Design news (Saturday 2026-09-05): Vendors sold fidelity as design skill this week — OpenAI GPT-6 Astra (Sept 3) markets template adherence and business-standard slide decks ($10/$50 per M tokens); Figma shipped agent-authored generative plugins/shaders (Sept 1), opacity as scoped number variable (Sept 3), Enterprise-managed MCP authorization GA (Aug 24), and agent chat panel in its own window (Aug 26). Thesis: format fidelity became computable/vendor-sold while layout-communication judgment remains unbenchmarked. | deepseek-v4-flash | 9.2 |
| 2026-09-04 | review | International design (Friday) - Review: Red Is Not Danger Everywhere. Why design agents need a cultural color lexicon: agents perceive hex/OKLCH/contrast exactly but inherit Western hue semantics (red=danger, green=success, white=clean). WCAG 1.4.3 contrast math is universal (4.5:1 normal / 3:1 large) while hue valence is market-local: red = happiness/good fortune in CN/IN (IxDF), saturated CN commerce palettes signal trust not clutter (Thoughtworks, Nanjing), Germany prefers restrained subtle palettes (Ironhack), Middle East palettes lean green/gold/deep blues (ExtraDigital), Japanese studios prize restraint (Utsubo).plan: deepseek-v4-flash | deepseek-v4-flash | 8.1 |
| 2026-09-03 | review | AI design models (Thursday) — Review: When to Fine-Tune a Design Agent. A routing rule for choosing between hosted prompting, managed fine-tuning, and open-weight deployment when teaching an AI agent to work inside a specific design system (task specificity, data availability, latency/cost, IP constraints as the decision axes).plan: mimo-v2.5-pro | deepseek-v4-flash | 8.9 |
| 2026-09-02 | review | Design trends (Wednesday) — Review: Can AI Agents Perceive Calm? 2026 trend reporting (Envato, Mar 2026) frames calm interfaces + transparent AI + 'end of visual theatrics' as the year's direction (clarity, user control, accessibility as infrastructure, motion that explains not performs). Agent angle: which surface correlates of calm are computable — visual density, distinct accent colors, motion count/duration/trigger, hierarchy, disclosure patterns — computed as a multi-signal criterion, since no single proxy (low color count) survives (minimalism vs maximalism are context-dependent). Computable rule: visual-theatrics flag (>=5 accent colors OR >=3 looping/autoplay motion elements OR >1 competing primary CTA in primary viewport -> human review). Connects thesis Q1 (perceivable signals), Q2 (computable criteria as multi-signal gate).plan: mimo-v2.5-pro | glm-5.3-flash | 8.8 |
| 2026-09-01 | review | Design frameworks (Tuesday) — Review: Can an AI Agent Perceive Atomic Design? Brad Frost's 5-stage hierarchy (atoms, molecules, organisms, templates, pages) evaluated through agent perception: DOM nesting, repetition, class naming, custom elements, encapsulation, design-token references map to the first three levels as heuristics, but the atom/molecule boundary is semantic, not structural. Computable rule: atomic-hierarchy-detection (confidence >= 0.70 gate, else 'unresolved semantic boundary'). Connects thesis Q1 (what can agents perceive) and Q2 (what criteria can they use).plan: mimo-v2.5-pro | glm-5.3-flash | 8.4 |
| 2026-08-31 | comparison | Design tools (Monday) — Comparison: Penpot's open .penpot file format (ZIP + JSON + SVG, manifest.json + tokens.json) vs Figma's proprietary .fig binary, scored on agent-computable criteria: file openness, SVG access, raw geometry inspection, token/theme analysis, contrast computation, automation surface, ecosystem maturity. Computable rule: open-structured-file-format gate.plan: mimo-v2.5-pro | glm-5.3-flash | 8.6 |
| 2026-08-30 | research | Design research (Sunday) — corpus fluid-literacy scan: do the sites agents learn from teach component-level responsiveness? Automated CSS-technique scan of 10 design-inspiration sources (2026-08-30): 6/10 CSS custom properties, 5/10 dark mode, 2/10 clamp(), 0/10 container queries. Computable rule: corpus-fluid-literacy gate.plan: mimo-v2.5-pro | glm-5.3-flash | 8.4 |
| 2026-08-29 | news | Design news (Saturday) — three answers to 'how should AI agents design?': Builder.io Agent Native Design (open-source artifact-native Figma alternative, Aug 6 2026), Figma AI credit add-ons expanded same-cost + Weave/agent to GA (Aug 25 2026), Claude Design prompt-to-prototype context (Apr 17 2026). Computable rule: vendor-domain inline-source lint.plan: mimo-v2.5-pro | poolside/laguna-s-2.1 | 8.3 |
| 2026-08-28 | standard | International design — script-aware typography metrics: why Western-trained design agents clip Devanagari matras, collide Arabic diacritics, and never wrap CJK, and the computable per-script rules that fix itplan: mimo-v2.5-pro | glm-5.3-flash | 9.2 |
| 2026-08-26 | review | Design Trends — Neobrutalism: The Most Agent-Computable Aesthetic of 2026 (thick borders, hard offset shadows, flat fills = literal CSS values an agent can grade with arithmetic)plan: mimo-v2.5-pro | deepseek-v4-flash | 8.2 |
| 2026-08-25 | test | Design Frameworks — Gestalt Principles Compile to Geometry: what an agent can actually compute (proximity/similarity/alignment/common-region PASS, figure-ground/closure FAIL)plan: deepseek-v4-flash | mimo-v2.5 | 8.0 |
| 2026-08-24 | standard | Maze (maze.co) — user testing / usability research platform as a feedback-signal generator for AI design agentsplan: mimo-v2.5-pro | mimo-v2.5 | 9.5 |
| 2026-08-24 | standard | Figma — industry standard UI design tool: what AI agents learn from auto layout, variables, components, and the official MCP serverplan: deepseek-v4-flash | poolside/laguna-s-2.1 | 8.8 |
| 2026-08-23 | standard | Design research — OKLCH perceptual color space: agents parse color as hex/RGB (device-centric, non-uniform), OKLCH is baseline since May 2023 and makes lightness a hue-agnostic computable rule (L-delta >= 0.45), while APCA is NOT the WCAG3 contrast algorithm (pulled July 2023 draft; 'yet to be determined' as of April 2026) — agents need provenance metadata on every rule.plan: mimo-v2.5-pro | deepseek-v4-flash | 8.5 |
| 2026-08-22 | standard | Design news — Anthropic's nine MCP connectors (April 28, 2026) plug Claude into Adobe Creative Cloud, Blender, Ableton Live, Autodesk Fusion, Splice, SketchUp, Affinity by Canva, Resolume Arena/Wire; bidirectional vs read-only tiers define agent-readiness; MCP tool-readiness heuristic for the scan.plan: mimo-v2.5-pro | deepseek-v4-flash | 8.4 |
| 2026-08-21 | standard | International design — text/information density as an agent-perceivable signal across markets (Japan 'text avalanche', China busy=credibility, Germany functional clarity, India multilingual density, Arabic RTL ornate density); density-context classifier rule for agents.plan: mimo-v2.5-pro | mimo-v2.5 | 8.6 |
| 2026-08-19 | — | Design trends — 3D/WebGL as functional interface layer in 2026; canvas/WebGL blind spot for DOM-inspecting AI agents (single <canvas> node, no scene graph, no computed styles for scene); AGENT_INSPECTABILITY_SCORE 0-4 semantic coverage ruleplan: mimo-v2.5-pro | poolside/laguna-s-2.1 | 9.0 |
| 2026-08-18 | — | Design frameworks — design tokens as the agent-native design framework; token coverage ratio as a computable proxy for agent-readability (DTCG spec 2025.10 stable, Material 3 token hierarchy, Style Dictionary, Tokens Studio)plan: mimo-v2.5-pro | mimo-v2.5 | 8.8 |
| 2026-08-17 | content | What an AI design agent can perceive in Rive's state-machine-based animation (.riv named states vs timeline keyframes)plan: mimo-v2.5-pro | deepseek-v4-flash | 8.1 |
| 2026-08-17 | — | Color Oracle (color blindness simulator) review — color perception as a computable signal; simulate-then-measure rule (WCAG contrast scored on Viénot/Brettel/Mollon-simulated colors)plan: deepseek-v4-flash | deepseek-v4-flash | 8.4 |
| 2026-08-16 | content | Design research: Can design quality be measured? Birkhoff 1933 M=O/C -> Ngo et al. 2003 (14 computable layout measures); Ivory & Hearst 2001 automation precedent; Reinecke CHI 2013 first-impression proxy trap; Cowan 2001 corrects Miller 7+-2 to 4 chunks; Nielsen heuristics are qualitative rules of thumb. What agents can safely compute as reward signals.plan: deepseek-v4-flash | deepseek-v4-flash | 8.7 |
| 2026-08-15 | content | Design news: Figma Agent — launched May 20, 2026 on the collaborative canvas; open beta at Config 2026 (June 24) with custom skills, web search, MCP connectors, attachments, generative plugins, shaders, and Motion. What this tells agents about the tools they need: perception, verification, design-system constraint, code handoff.plan: mimo-v2.5-pro | deepseek-v4-flash | 8.4 |
| 2026-08-14 | content | International design: what an AI design agent can perceive about RTL (Arabic) web design — direction as a computable design signal (BiDi, dir/direction/unicode-bidi, logical vs physical CSS, 20% physical-property gate)plan: mimo-v2.5-pro | poolside/laguna-s-2.1 | 7.8 |
| 2026-08-13 | standard | AI design models: mapping 15 design models by perception modality and output mediumplan: deepseek-v4-flash | deepseek-v4-flash | 8.1 |
| 2026-08-12 | standard | Design Trends: resolution bias makes AI agents default to minimalism in a maximalist 2026plan: deepseek-v4-flash | mimo-v2.5 | 8.1 |
| 2026-08-11 | standard | Design Frameworks: Fitts's + Hick's Laws as agent-computable arithmeticplan: deepseek-v4-flash | deepseek-v4-flash | 9.7 |
| 2026-08-10 | standard | Design Tools: which design-to-code builder (v0.dev, Webflow, Framer, Relume) produces output an AI agent can parse and verifyplan: deepseek-v4-flash | deepseek-v4-flash | 8.8 |
| 2026-08-10 | standard | Axe (Deque Systems) — Automated Accessibility CI Engineplan: mimo-v2.5-pro | deepseek-v4-flash | 8.5 |
| 2026-08-08 | standard | Design News: 94.8% web accessibility failure rate and 5,114 ADA lawsuits as an agent-computable design signal | poolside/laguna-s-2.1 | 9.1 |
| 2026-08-07 | standard | International Design: why a Western-trained design agent misjudges Japan, China, and Arabic web conventions on computable signals | deepseek-v4-flash | 8.9 |
| 2026-08-05 | standard | Design Trends: the polarized 2026 web — what an AI agent can perceive | deepseek-v4-flash | 9.3 |
| 2026-08-04 | standard | Design Frameworks: which frameworks survive translation into computable agent rules | deepseek-v4-flash | 9.0 |
| 2026-08-03 | standard | Design Tools: 2026 design-tool market read by an AI agent | deepseek-v4-flash | 9.3 |
| 2026-08-02 | standard | Design Research: container-query gap in the design-inspiration corpus | mimo-v2.5 | 9.2 |
| 2026-08-01 | standard | Design News: Claude Design launch | mimo-v2.5 | 8.2 |
| 2026-07-31 | standard | Design Tool Landscape Survey | deepseek-v4-flash | 9.0 |
| 2026-07-28 | standard | 8-Point Grid Spacing | — | — |
| 2026-07-24 | standard | International Design | — | — |
Model pool
8 models · by provider| Model | Role | Status |
|---|---|---|
| deepseek | ||
| deepseek-v4-flash | builder | untested |
| openrouter | ||
| poolside/laguna-s-2.1 | builder | untested |
| qwen3.7-flash | judge | untested |
| openai/gpt-5.6-luna | research | untested |
| xiaomi | ||
| mimo-v2.5 | builder | untested |
| mimo-v2.5-pro | planner | untested |
| zai | ||
| glm-5.2 | qa | untested |
| glm-5.3-flash | builder | untested |