Job description
AI Infrastructure Engineer Own Your Impact. At Propio , we don't believe careers happen to people. We believe people create them.
Here, you're trusted to make decisions, challenge assumptions, drive innovation, and shape outcomes. Your success is not limited by hierarchy or tenure. It's fueled by your ambition, your curiosity, and your willingness to own your impact.
If you're looking for a role where you can simply maintain the status quo, this probably isn't it, but if you're looking for a place where your ideas matter, your growth is accelerated, and your work creates meaningful impact across the world, we'd love to talk. Why Propio? Every day, communication changes lives.
A patient receives care they otherwise couldn't access. A family gains critical information. A business connects with a customer.
A community becomes more inclusive. These moments happen because barriers are removed. And behind those moments are Propio team members who show up every day to solve problems, innovate, and build the future.
This isn't just work. This is world impact. As an AI Infrastructure Engineer, you will primarily design, build, and operate our AWS-based, GPU-accelerated inference and streaming serving platform, optimizing it for low latency, high concurrency, reliability, and cost.
You will also support the research team’s training environment, further the AI team’s ML/LLMOps capabilities, and enable the development of edge AI deployments. You'll be empowered to: Take ownership of important initiatives and outcomes. Drive meaningful business results.
Influence decisions and contribute new ideas. Partner with talented, high-performing team members. Challenge yourself through continuous learning and growth.
Help shape the future of a rapidly growing organization What You'll Own Real-time inference and serving. Design, deploy, and operate low-latency streaming inference systems on AWS for LLM, ASR, TTS, and multimodal models. Optimize time to first token/audio, p95/p99 end-to-end latency, real-time factor, throughput, concurrency, GPU utilization, and cost per stream using technologies such as vLLM, SGLang, TensorRT-LLM, Triton, TensorRT.
Training environment. Support and evolve the research team’s AWS-based training environment, including reproducible containers, GPU job scheduling, distributed job execution, checkpoint and resume capabilities, experiment tracking, model and data artifact access, and researcher self-service. Enable workloads using FSDP, DeepSpeed, Megatron-LM, Slurm, EKS, or SageMaker.
ML/LLMOps. Build model and artifact registries, lineage and versioning, automated evaluation gates, deployment pipelines, shadow and canary releases, rollback workflows, runtime and configuration management, and production observability Edge AI infrastructure. Partner with researchers and device and embedded teams to build repeatable model optimization, packaging, validation, and deployment workflows for resource-constrained edge targets.
Support model export and compilation, post-training quantization, runtime integration. Reliability and security. Own capacity planning, production readiness, incident response, disaster recovery, and the secure operation of the AI platform.
Apply least-privilege IAM, KMS encryption, private networking, secrets management. What Makes Someone Successful Here The most successful people at Propio aren't necessarily the ones with the longest resumes. They're the people who: Take ownership instead of waiting for direction.
Embrace challenges as opportunities to grow. Continuously seek better ways of working. Turn ideas into action.
Hold themselves and others accountable to high standards. Are driven by making a measureable impact.
