Notes
August 2026
Continual Harness: Agents improve their own tools, skills, and memory mid-run
August 2026
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
July 2026
DeepSeek - V4 Notes
July 2026
Purified On-policy Self-distillation Paper Notes
July 2026
Reinforcement Learning via Self-Distillation Paper Notes
July 2026
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
June 2026
Introduction to Reinforcement Learning Notes
June 2026
RL Foundations and PPO Notes
June 2026
On-Policy Self-Distillation for Large Language Models
June 2026
Self-Distillation Enables Continual Learning
June 2026
Adapting the Interface, Not the Model
June 2026
AGENTS.md Paper Notes
June 2026
CaMeL: Computer Use Agents
June 2026
DeepDive Paper Notes
June 2026
DeepSeekMath Paper Notes
June 2026
DeepSeek-R1 Paper Notes
June 2026
Natural-Language Agent Harnesses
June 2026
Recursive Language Models
June 2026
SFR-DeepResearch Notes
October 2025
Process Reward Model in RL Training
July 2025
Model parallelization techniques