swe bench pro qwen refined and corrected task set Benchmark: Scores and Sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: coding · weight 1 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 Qwen3.8 Max Alibaba / Qwen 67.7%Official model cards via models.devLab-reported; metric resolved; transcribed by MIT models.dev catalog; not independently evaluated [variant] Claude Code; Qwen refined and corrected task set Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
67.7 85% confidence 85 percent, High
2 Qwen3.8 Flash Next Alibaba / Qwen 62.5%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] Claude Code; Qwen refined and corrected task set Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
62.5 21% confidence 21 percent, Low

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the coding pillar of the SI Score.