Location
Remote, Remote, Philippines
Posted
July 25, 2026
Job Description
About the Role
As a Senior AI/LLM Engineer, you will lead our efforts to train, align, and optimize large language models. You will own the full post-training pipeline from supervised fine-tuning through reward modeling and RL optimization, while also ensuring models run efficiently in production. This is a role that bridges alignment research and systems engineering.
What You’ll Own
- Own and drive the full RLHF pipeline: data collection, reward model training, and RL fine-tuning using PPO, DPO, GRPO, and RLAIF
- Design and run Supervised Fine-Tuning (SFT) pipelines on open-weight models (LLaMA, Mistral, Qwen) as the foundation for RLHF
- Build and train reward models that accurately capture human preferences from annotation data
- Design human feedback collection pipelines: labeling rubrics, annotator calibration, and preference dataset curation
- Implement Constitutional AI and RLAI...