Now
What I am working on, thinking about, and reading. Updated monthly.
Last updated
- Work
- Post-training for anime-domain video generation at Tencent Hunyuan. The current focus is a unified generation architecture, meaning one model that handles text-to-video, image-to-video, multi-reference generation, and editing, along with inference acceleration and distillation.
- Side project
- A pipeline that turns papers into narrated videos. Feed it an arXiv paper, get back a video with narration, shot planning, and animation. It started because the hard part of reading a paper is rarely the math. It is building intuition for the method as a whole, and video is much better at that than text.
- Thinking about
- Where the evaluation signal for preference alignment in video generation should come from. Human annotation is slow and expensive, automatic metrics often disagree with human judgment, and online behavioral data is entangled with factors unrelated to generation quality. I do not have a good answer yet.
- Reading
- Reinforcement learning applied to generative models, and work related to world models. Also catching up on some non-technical reading.
The idea for this kind of page comes from the /now page movement started by Derek Sivers.