Open Role
Inference Systems Engineer
Singapore
Build high-performance inference infrastructure that puts embodied AI models on real robots.
Responsibilities
- Design, build, and maintain inference systems for Vision-Language-Action (VLA), manipulation, and whole-body control models deployed on robot hardware.
- Optimize model serving for low-latency, closed-loop robot execution across datacenter GPUs, edge compute, and on-robot inference platforms.
- Implement performance optimizations including graph compilation, kernel fusion, quantization, batching strategies, and memory management for production workloads.
- Develop runtime bridges between research models and robot software stacks, ensuring reliable tensor I/O, synchronization, and fault tolerance.
- Build monitoring, profiling, and observability tooling to track latency, throughput, GPU utilization, and deployment health across robot fleets.
- Collaborate with research and robotics teams to support model compression, distillation, and deployment-ready export pipelines.
- Design rollout and versioning systems that enable safe A/B testing, canary releases, and rollback of model updates in real-world environments.
- Contribute to internal standards for inference APIs, benchmarking, and reproducible performance evaluation across embodiments and tasks.
Qualifications
- A deep and genuine passion for Embodied AI is essential.
- Bachelor's degree or above in Computer Science, Electrical Engineering, or related fields. Advanced degree is a plus.
- 3+ years of experience in systems engineering, ML infrastructure, or high-performance computing.
- Strong programming skills in C++ and Python, with solid understanding of concurrency, memory, and performance trade-offs.
- Hands-on experience with GPU programming and optimization (CUDA, TensorRT, Triton, ONNX Runtime, or similar).
- Experience deploying deep learning models in production, including serving frameworks, containerization, and distributed systems.
- Familiarity with PyTorch model export, compilation workflows, and debugging numerical or runtime issues in deployed models.
- Understanding of real-time and near-real-time system constraints relevant to robotics and embodied AI applications.
- Experience with robot middleware (ROS/ROS2), edge deployment, or on-device inference is highly preferred.