Cursor’s proprietary Composer 2.5 model has been released, and its benchmarks are nearly indistinguishable from Anthropic’s Claude Opus 4.7. On Terminal-Bench 2.0, Composer 2.5 scored 69.3% versus Opus 4.7’s 69.4%. In SWE-Bench Multilingual, the gap was similarly narrow: 79.8% for Composer against 80.5% for Opus. On CursorBench v3.1, Composer 2.5 also performed competitively. The company has essentially built a model that goes head-to-head with Anthropic’s offering.
3mo
Cursor’s proprietary Composer 2.5 model has been released, and its benchmarks are nearly indistinguishable from Anthropic’s Claude Opus 4.7. On Terminal-Bench 2.0, Composer 2.5 scored 69.3% versus Opus 4.7’s 69.4%. In SWE-Bench Multilingual, the gap was similarly narrow: 79.8% for Composer against 80.5% for Opus. On CursorBench v3.1, Composer 2.5 also performed competitively. The company has essentially built a model that goes head-to-head with Anthropic’s offering.
3mo
아직 댓글이 없습니다. 첫 댓글을 남겨보세요!
댓글
아직 댓글이 없습니다. 첫 댓글을 남겨보세요!