Notes

August 2026

Continual Harness: Agents improve their own tools, skills, and memory mid-run

August 2026

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

July 2026

DeepSeek - V4 Notes

July 2026

Purified On-policy Self-distillation Paper Notes

July 2026

Reinforcement Learning via Self-Distillation Paper Notes

July 2026

On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

June 2026

Introduction to Reinforcement Learning Notes

June 2026

RL Foundations and PPO Notes

June 2026

On-Policy Self-Distillation for Large Language Models

June 2026

Self-Distillation Enables Continual Learning

June 2026

Adapting the Interface, Not the Model

June 2026

AGENTS.md Paper Notes

June 2026

CaMeL: Computer Use Agents

June 2026

DeepDive Paper Notes

June 2026

DeepSeekMath Paper Notes

June 2026

DeepSeek-R1 Paper Notes

June 2026

Natural-Language Agent Harnesses

June 2026

Recursive Language Models

June 2026

SFR-DeepResearch Notes

October 2025

Process Reward Model in RL Training

July 2025

Model parallelization techniques