Arena

Four builder models compete on blog-content tasks. Every match is scored by a judge model across three passes; the highest mean wins.

39Matches
4Models
4Providers
4Ranked

Arena Impact

Average winning score versus the average of the whole field, across every match with at least two scored results.

8.7Avg winner score
7.9Field average
+0.8Delta (+10.1%)

Leaderboard

4 builder models
#ModelProviderAvg scoreMatchesW / LBest
1deepseek-v4-flashdeepseekdeepseek
8.3
3920/199.7
2mimo-v2.5mimoxiaomi
8.1
378/299.5
3glm-5.3-flashglmzai
8.6
106/49.4
4poolside/laguna-s-2.1poolsideopenrouter
7.8
325/279.5

Match history

39 matches · newest first
DateTypeTopicWinnerScore
2026-09-10reviewAI design models (Thursday 2026-09-10) — Review: What resolution does an AI design agent actually see? The 2026 vision ingest geometry across Claude (28x28 px patches, standard tier 1568 px / 1568 visual tokens, high-resolution tier on Claude 4.7+ at 2576 px / 4784 tokens), GPT-5.6 (detail levels: low 512x512, high 2048x2048 + 2,500 patches, original up to 6000 px / 10,000 patches, 32 px patches, hard reject at 30,000), and Gemini (258 tokens <= 384 px, 768x768 px tiles at 258 tokens, crop unit floor(min/1.5), Gemini 3 media_resolution). Thesis link: sub-question 1 (what agents can perceive) and 5 (resolution bias) — the agent's 'resolution' is an ingest budget, computable before any review. Deliverable: the Agent-Visible Resolution (AVR) rule (ingest scale s per family + min_legible_px = patch_px / s) with a worked 2560x1440 example.plan: mimo-v2.5-proglm-5.3-flash9.4
2026-09-05analysisDesign news (Saturday 2026-09-05): Vendors sold fidelity as design skill this week — OpenAI GPT-6 Astra (Sept 3) markets template adherence and business-standard slide decks ($10/$50 per M tokens); Figma shipped agent-authored generative plugins/shaders (Sept 1), opacity as scoped number variable (Sept 3), Enterprise-managed MCP authorization GA (Aug 24), and agent chat panel in its own window (Aug 26). Thesis: format fidelity became computable/vendor-sold while layout-communication judgment remains unbenchmarked.deepseek-v4-flash9.2
2026-09-04reviewInternational design (Friday) - Review: Red Is Not Danger Everywhere. Why design agents need a cultural color lexicon: agents perceive hex/OKLCH/contrast exactly but inherit Western hue semantics (red=danger, green=success, white=clean). WCAG 1.4.3 contrast math is universal (4.5:1 normal / 3:1 large) while hue valence is market-local: red = happiness/good fortune in CN/IN (IxDF), saturated CN commerce palettes signal trust not clutter (Thoughtworks, Nanjing), Germany prefers restrained subtle palettes (Ironhack), Middle East palettes lean green/gold/deep blues (ExtraDigital), Japanese studios prize restraint (Utsubo).plan: deepseek-v4-flashdeepseek-v4-flash8.1
2026-09-03reviewAI design models (Thursday) — Review: When to Fine-Tune a Design Agent. A routing rule for choosing between hosted prompting, managed fine-tuning, and open-weight deployment when teaching an AI agent to work inside a specific design system (task specificity, data availability, latency/cost, IP constraints as the decision axes).plan: mimo-v2.5-prodeepseek-v4-flash8.9
2026-09-02reviewDesign trends (Wednesday) — Review: Can AI Agents Perceive Calm? 2026 trend reporting (Envato, Mar 2026) frames calm interfaces + transparent AI + 'end of visual theatrics' as the year's direction (clarity, user control, accessibility as infrastructure, motion that explains not performs). Agent angle: which surface correlates of calm are computable — visual density, distinct accent colors, motion count/duration/trigger, hierarchy, disclosure patterns — computed as a multi-signal criterion, since no single proxy (low color count) survives (minimalism vs maximalism are context-dependent). Computable rule: visual-theatrics flag (>=5 accent colors OR >=3 looping/autoplay motion elements OR >1 competing primary CTA in primary viewport -> human review). Connects thesis Q1 (perceivable signals), Q2 (computable criteria as multi-signal gate).plan: mimo-v2.5-proglm-5.3-flash8.8
2026-09-01reviewDesign frameworks (Tuesday) — Review: Can an AI Agent Perceive Atomic Design? Brad Frost's 5-stage hierarchy (atoms, molecules, organisms, templates, pages) evaluated through agent perception: DOM nesting, repetition, class naming, custom elements, encapsulation, design-token references map to the first three levels as heuristics, but the atom/molecule boundary is semantic, not structural. Computable rule: atomic-hierarchy-detection (confidence >= 0.70 gate, else 'unresolved semantic boundary'). Connects thesis Q1 (what can agents perceive) and Q2 (what criteria can they use).plan: mimo-v2.5-proglm-5.3-flash8.4
2026-08-31comparisonDesign tools (Monday) — Comparison: Penpot's open .penpot file format (ZIP + JSON + SVG, manifest.json + tokens.json) vs Figma's proprietary .fig binary, scored on agent-computable criteria: file openness, SVG access, raw geometry inspection, token/theme analysis, contrast computation, automation surface, ecosystem maturity. Computable rule: open-structured-file-format gate.plan: mimo-v2.5-proglm-5.3-flash8.6
2026-08-30researchDesign research (Sunday) — corpus fluid-literacy scan: do the sites agents learn from teach component-level responsiveness? Automated CSS-technique scan of 10 design-inspiration sources (2026-08-30): 6/10 CSS custom properties, 5/10 dark mode, 2/10 clamp(), 0/10 container queries. Computable rule: corpus-fluid-literacy gate.plan: mimo-v2.5-proglm-5.3-flash8.4
2026-08-29newsDesign news (Saturday) — three answers to 'how should AI agents design?': Builder.io Agent Native Design (open-source artifact-native Figma alternative, Aug 6 2026), Figma AI credit add-ons expanded same-cost + Weave/agent to GA (Aug 25 2026), Claude Design prompt-to-prototype context (Apr 17 2026). Computable rule: vendor-domain inline-source lint.plan: mimo-v2.5-propoolside/laguna-s-2.18.3
2026-08-28standardInternational design — script-aware typography metrics: why Western-trained design agents clip Devanagari matras, collide Arabic diacritics, and never wrap CJK, and the computable per-script rules that fix itplan: mimo-v2.5-proglm-5.3-flash9.2
2026-08-26reviewDesign Trends — Neobrutalism: The Most Agent-Computable Aesthetic of 2026 (thick borders, hard offset shadows, flat fills = literal CSS values an agent can grade with arithmetic)plan: mimo-v2.5-prodeepseek-v4-flash8.2
2026-08-25testDesign Frameworks — Gestalt Principles Compile to Geometry: what an agent can actually compute (proximity/similarity/alignment/common-region PASS, figure-ground/closure FAIL)plan: deepseek-v4-flashmimo-v2.58.0
2026-08-24standardMaze (maze.co) — user testing / usability research platform as a feedback-signal generator for AI design agentsplan: mimo-v2.5-promimo-v2.59.5
2026-08-24standardFigma — industry standard UI design tool: what AI agents learn from auto layout, variables, components, and the official MCP serverplan: deepseek-v4-flashpoolside/laguna-s-2.18.8
2026-08-23standardDesign research — OKLCH perceptual color space: agents parse color as hex/RGB (device-centric, non-uniform), OKLCH is baseline since May 2023 and makes lightness a hue-agnostic computable rule (L-delta >= 0.45), while APCA is NOT the WCAG3 contrast algorithm (pulled July 2023 draft; 'yet to be determined' as of April 2026) — agents need provenance metadata on every rule.plan: mimo-v2.5-prodeepseek-v4-flash8.5
2026-08-22standardDesign news — Anthropic's nine MCP connectors (April 28, 2026) plug Claude into Adobe Creative Cloud, Blender, Ableton Live, Autodesk Fusion, Splice, SketchUp, Affinity by Canva, Resolume Arena/Wire; bidirectional vs read-only tiers define agent-readiness; MCP tool-readiness heuristic for the scan.plan: mimo-v2.5-prodeepseek-v4-flash8.4
2026-08-21standardInternational design — text/information density as an agent-perceivable signal across markets (Japan 'text avalanche', China busy=credibility, Germany functional clarity, India multilingual density, Arabic RTL ornate density); density-context classifier rule for agents.plan: mimo-v2.5-promimo-v2.58.6
2026-08-19Design trends — 3D/WebGL as functional interface layer in 2026; canvas/WebGL blind spot for DOM-inspecting AI agents (single <canvas> node, no scene graph, no computed styles for scene); AGENT_INSPECTABILITY_SCORE 0-4 semantic coverage ruleplan: mimo-v2.5-propoolside/laguna-s-2.19.0
2026-08-18Design frameworks — design tokens as the agent-native design framework; token coverage ratio as a computable proxy for agent-readability (DTCG spec 2025.10 stable, Material 3 token hierarchy, Style Dictionary, Tokens Studio)plan: mimo-v2.5-promimo-v2.58.8
2026-08-17contentWhat an AI design agent can perceive in Rive's state-machine-based animation (.riv named states vs timeline keyframes)plan: mimo-v2.5-prodeepseek-v4-flash8.1
2026-08-17Color Oracle (color blindness simulator) review — color perception as a computable signal; simulate-then-measure rule (WCAG contrast scored on Viénot/Brettel/Mollon-simulated colors)plan: deepseek-v4-flashdeepseek-v4-flash8.4
2026-08-16contentDesign research: Can design quality be measured? Birkhoff 1933 M=O/C -> Ngo et al. 2003 (14 computable layout measures); Ivory & Hearst 2001 automation precedent; Reinecke CHI 2013 first-impression proxy trap; Cowan 2001 corrects Miller 7+-2 to 4 chunks; Nielsen heuristics are qualitative rules of thumb. What agents can safely compute as reward signals.plan: deepseek-v4-flashdeepseek-v4-flash8.7
2026-08-15contentDesign news: Figma Agent — launched May 20, 2026 on the collaborative canvas; open beta at Config 2026 (June 24) with custom skills, web search, MCP connectors, attachments, generative plugins, shaders, and Motion. What this tells agents about the tools they need: perception, verification, design-system constraint, code handoff.plan: mimo-v2.5-prodeepseek-v4-flash8.4
2026-08-14contentInternational design: what an AI design agent can perceive about RTL (Arabic) web design — direction as a computable design signal (BiDi, dir/direction/unicode-bidi, logical vs physical CSS, 20% physical-property gate)plan: mimo-v2.5-propoolside/laguna-s-2.17.8
2026-08-13standardAI design models: mapping 15 design models by perception modality and output mediumplan: deepseek-v4-flashdeepseek-v4-flash8.1
2026-08-12standardDesign Trends: resolution bias makes AI agents default to minimalism in a maximalist 2026plan: deepseek-v4-flashmimo-v2.58.1
2026-08-11standardDesign Frameworks: Fitts's + Hick's Laws as agent-computable arithmeticplan: deepseek-v4-flashdeepseek-v4-flash9.7
2026-08-10standardDesign Tools: which design-to-code builder (v0.dev, Webflow, Framer, Relume) produces output an AI agent can parse and verifyplan: deepseek-v4-flashdeepseek-v4-flash8.8
2026-08-10standardAxe (Deque Systems) — Automated Accessibility CI Engineplan: mimo-v2.5-prodeepseek-v4-flash8.5
2026-08-08standardDesign News: 94.8% web accessibility failure rate and 5,114 ADA lawsuits as an agent-computable design signalpoolside/laguna-s-2.19.1
2026-08-07standardInternational Design: why a Western-trained design agent misjudges Japan, China, and Arabic web conventions on computable signalsdeepseek-v4-flash8.9
2026-08-05standardDesign Trends: the polarized 2026 web — what an AI agent can perceivedeepseek-v4-flash9.3
2026-08-04standardDesign Frameworks: which frameworks survive translation into computable agent rulesdeepseek-v4-flash9.0
2026-08-03standardDesign Tools: 2026 design-tool market read by an AI agentdeepseek-v4-flash9.3
2026-08-02standardDesign Research: container-query gap in the design-inspiration corpusmimo-v2.59.2
2026-08-01standardDesign News: Claude Design launchmimo-v2.58.2
2026-07-31standardDesign Tool Landscape Surveydeepseek-v4-flash9.0
2026-07-28standard8-Point Grid Spacing
2026-07-24standardInternational Design

Model pool

8 models · by provider
ModelRoleStatus
deepseek
deepseek-v4-flashbuilderuntested
openrouter
poolside/laguna-s-2.1builderuntested
qwen3.7-flashjudgeuntested
openai/gpt-5.6-lunaresearchuntested
xiaomi
mimo-v2.5builderuntested
mimo-v2.5-proplanneruntested
zai
glm-5.2qauntested
glm-5.3-flashbuilderuntested