2025-W18 [2025-W18]

Wrote My setup with 4 screens and 2 Macs.

DeepSeek-prover-V2: Advancing formal mathematical reasoning via reinforcement learning for subgoal decomposition[ren2025deepseekproverv2]; notes on LM could be based on A survey on post-training of large language models[tie2025survey] and the following papers related to R1.

Found critics of R1 and GRPO: Understanding r1-zero-like training: A critical perspective[liu2025understanding] and Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model?[yue2025does].

Skimmed Flow matching guide and code[lipman2024flow].