Hello! I’m Thanawat or Sky. I’m a PhD student at The University of Tokyo under the supervision of Prof. Takashi Ishida and Prof. Masashi Sugiyama.
Outside of research
- I love football (soccer), both watching and playing. I support Arsenal, and my all-time favorite player is Mesut Özil. I also play volleyball and badminton, though not as well as football. I often do calisthenics workouts, but I’m still a beginner.
- I love listening to sad-tone songs. Some of my favorites are Everything by The Black Skirts, あの夢をなぞって (Ballade Ver.) by YOASOBI, About You by The 1975, เจ็บจนไม่เข้าใจ by PORTRAIT, and ถ้าเธอ by STAMP & Violette Wautier.
- I love listening to history stories.
- I absolutely love cooking, both for myself and for others — it might be my favorite hobby.
- I love dogs.
My research interests are unintended/unaligned behaviors in LLMs and statistical safeguarding for LLMs. For example:
- In NeurIPS 2026, we detect coding agents gaming test suites with CapCode, which designs the Bayes error of a coding benchmark: injecting controlled randomness sets a known cap on the best possible score (unit test pass rate), strictly below 100%, and prevent hacking rewards in RL post-training with CapReward, which rewards as usual when the score is below the cap but penalizes it once the score goes beyond the cap.
- In ICML 2026, we introduce CapBencher, which extends the same cap to benchmarks without unit tests, turning any score above it into a built-in alarm that flags test-set overfitting and benchmark data being contaminated.
- In TMLR 2025, we study importance weighting for keeping LLMs aligned under deployment distribution shift where training and test data come from different objectives (e.g., training data optimized for helpfulness but test-time deployment also requiring harmlessness).