Comparing visual spec-to-app model scores over time (updated periodically)
| Rank | Model | Score | Delta | Notes |
|---|---|---|---|---|
| 1 | GPT-4o (2024-08) | 92 | +4 | Strong spec compliance, handles complex UIs |
| 2 | Claude 3.5 Sonnet | 88 | +2 | Excellent layout accuracy, slower generation |
| 3 | Gemini 1.5 Pro | 85 | +3 | Good for simple forms, struggles with animations |
| 4 | Claude 3 Opus | 83 | -1 | Baseline benchmark from Q2 2024 |
| 5 | GPT-4 Turbo | 80 | -2 | Still solid for basic spec-to-app tasks |
| 6 | Llama 3.1 405B | 74 | +5 | Open-source leader, improving rapidly |
| 7 | Mistral Large 2 | 70 | +1 | Good for minimal UIs, lacks image support |
| 8 | Gemma 2 27B | 65 | +3 | Promising small model for mobile UIs |
Copy the link to share this benchmark board: