- AI, But Simple
- Posts
- RL Training Inside Harnesses, Stolen Reasoning Traces, Judging Agents
RL Training Inside Harnesses, Stolen Reasoning Traces, Judging Agents
AI, But Simple Issue #118

RL Training Inside Harnesses, Stolen Reasoning Traces, Judging Agents
AI, But Simple Issue #118
This week, AI research has landed on training agents to maximize harness scores, research on stealing reasoning traces from models, and a focus on benchmarking the reliability of LLM-as-a-judge.
We’ll get into the developments of these fields, get into how their underlying mechanisms work, cover some benchmark results, and key research movements in the topics to keep you updated in no time.