Turing × Anthropic
CS Expert evaluating and improving Anthropic Claude Opus models (4.5 → 4.8 / Fable) via Turing.
Computer Science Expert at Turing, collaborating on Anthropic Claude Opus frontier models—including Opus 4.5, 4.6, and 4.8 (Fable).
What I worked on
The mandate was simple and hard: find every place the model fails, then make it better.
Work spanned high-stakes evaluation and improvement loops across:
- Data science — analytical reasoning, pipelines, statistical judgment, and tool-using workflows
- Machine learning — model understanding, experiment design, and ML problem-solving
- Research-paper tasks — reading, critiquing, and producing research-grade technical writing
- Multimedia models — multimodal failure analysis and quality improvements
Impact
Through systematic red-teaming and failure discovery, contributions fed into making Claude Opus generations more reliable on real expert tasks—not just benchmark scores.
Stack / focus: LLM evaluation, RL / preference data, multimodal reasoning, research workflows.