I am a Research Scientist working on reasoning and model adaptation through post-training, self-distillation, and reinforcement learning. My current research focuses on learning from long-horizon trajectories and developing methods for improving reasoning behavior through iterative post-training.
Previously, I was a Research Scientist at Meta, where I worked on internet-scale personalization and ranking systems. My work focused on learning from noisy implicit feedback, long-horizon value modeling, user representation, and debiasing large-scale decision systems.
I received my Ph.D. in Computer Science from the University of Pittsburgh, advised by Dr. Milos Hauskrecht. Across my research, a recurring theme has been how learning systems can adapt from sequential, heterogeneous, and imperfect feedback.
Current interests: reasoning and post-training; self-distillation and reinforcement learning; long-horizon learning; model adaptation; reward and preference modeling; evaluation and memory.