organizational research article
Evaluating frontier AI R&D capabilities of language model agents against human experts
METR · Published 2024-11-22
Evidence from this source
| Type | Evidence |
|---|---|
| Benchmark | RE-BenchSource detailsEnvironment descriptions; Results; What do we mean by time budget? |
| Paper / report | Evaluating frontier AI R&D capabilities of language model agents against human expertsSource detailsPrimary publication source |
Document tracking details
- Source ID
- src-rebench
- Source type
- Primary
- Original URL
- https://metr.org/blog/2024-11-22-evaluating-r-d-capabilities-of-llms/
- Last updated by publisher
- Not reported
- Access status
- available
- Retrieved
- 2026-09-26T04:47:27Z
Availability checks
- 2026-09-26T04:47:27Z · success: Original source inspected by Codex; extraction check, not experimental replication.