Job description
Applied Scientist/Research Engineer, LLM Training Data Own Your Impact. At Propio , we don't believe careers happen to people. We believe people create them.
Here, you're trusted to make decisions, challenge assumptions, drive innovation, and shape outcomes. Your success is not limited by hierarchy or tenure. It's fueled by your ambition, your curiosity, and your willingness to own your impact.
If you're looking for a role where you can simply maintain the status quo, this probably isn't it, but if you're looking for a place where your ideas matter, your growth is accelerated, and your work creates meaningful impact across the world, we'd love to talk. Why Propio? Every day, communication changes lives.
A patient receives care they otherwise couldn't access. A family gains critical information. A business connects with a customer.
A community becomes more inclusive. These moments happen because barriers are removed. And behind those moments are Propio team members who show up every day to solve problems, innovate, and build the future.
This isn't just work. This is world impact. As a Applied Scientist/Research Engineer, LLM Training Data, you'll have the opportunity to make a meaningful contribution to the continued growth and transformation of Propio.
You'll be empowered to: Take ownership of important initiatives and outcomes. Drive meaningful business results. Influence decisions and contribute new ideas.
Partner with talented, high-performing team members. Challenge yourself through continuous learning and growth. Help shape the future of a rapidly growing organization What You'll Own As an Applied Scientist/Research Engineer, LLM Training Data , you will own the data strategy, curation pipelines, annotation workflows, and evaluation datasets that power our multilingual AI systems.
This is a hands-on technical role for someone who understands how to manage the full AI data lifecycle, from acquisition, curation, annotation, and quality control to evaluation datasets and post-training data, to directly improve LLM performance. The ideal candidate can build scalable data pipelines, design high-quality annotation and QA processes, identify model failure modes, and close performance gaps through targeted data acquisition, curation, and synthetic data generation. Define the end-to-end data roadmap for multilingual and multimodal AI systems, including text, speech, translation, interpretation, low-resource languages, and agentic AI workflows.
Design and build dataset curation pipelines for training, post-training, and evaluation, including cleaning, deduplication, filtering, PII redaction, quality scoring, sampling, balancing, and versioning. Create annotation schemas, labeling guidelines, QA rubrics, golden datasets, and reviewer workflows for multilingual, speech, translation, and vision data. Build evaluation datasets and benchmarks, analyze model failure modes, and translate performance gaps into targeted data improvements.
Support post-training data workflows such as SFT, instruction tuning, preference data, RLHF/DPO-style data, reward model data, and synthetic data generation. Use modern annotation tools and AWS-based data infrastructure to scale secure, traceable, and compliant AI data workflows. What Makes Someone Successful Here The most successful people at Propio aren't necessarily the ones with the longest resumes.
They're the people who: Take ownership instead of waiting for direction. Embrace challenges as opportunities to grow. Continuously seek better ways of working.
Turn ideas into action. Hold themselves and others accountable to high standards. Are driven by making a measurable impact.
