• AI, But Simple
  • Posts
  • RL Training Inside Harnesses, Stolen Reasoning Traces, Judging Agents

RL Training Inside Harnesses, Stolen Reasoning Traces, Judging Agents

AI, But Simple Issue #118

RL Training Inside Harnesses, Stolen Reasoning Traces, Judging Agents

AI, But Simple Issue #118

This week, AI research has landed on training agents to maximize harness scores, research on stealing reasoning traces from models, and a focus on benchmarking the reliability of LLM-as-a-judge.

We’ll get into the developments of these fields, get into how their underlying mechanisms work, cover some benchmark results, and key research movements in the topics to keep you updated in no time.

Subscribe to keep reading

This content is free, but you must be subscribed to AI, But Simple to continue reading.

Already a subscriber?Sign in.Not now