Turing × Anthropic

CS Expert evaluating and improving Anthropic Claude Opus models (4.5 → 4.8 / Fable) via Turing.

Computer Science Expert at Turing, collaborating on Anthropic Claude Opus frontier models—including Opus 4.5, 4.6, and 4.8 (Fable).

What I worked on

The mandate was simple and hard: find every place the model fails, then make it better.

Work spanned high-stakes evaluation and improvement loops across:

  • Data science — analytical reasoning, pipelines, statistical judgment, and tool-using workflows
  • Machine learning — model understanding, experiment design, and ML problem-solving
  • Research-paper tasks — reading, critiquing, and producing research-grade technical writing
  • Multimedia models — multimodal failure analysis and quality improvements

Impact

Through systematic red-teaming and failure discovery, contributions fed into making Claude Opus generations more reliable on real expert tasks—not just benchmark scores.

Stack / focus: LLM evaluation, RL / preference data, multimodal reasoning, research workflows.