Inference Engineer
We are looking for a skilled and passionate Inference Engineer to design, optimize, and deploy high-performance AI inference pipelines for production environments. In this role, you will bridge the gap between machine learning models and high-throughput, low-latency production infrastructure. You will be responsible for optimizing state-of-the-art models (such as LLMs, Speech/Audio models, and Computer Vision architectures), reducing compute costs, and ensuring resilient, scalable serving on enterprise-grade hardware (GPUs/accelerators). If you have deep expertise in model quantization, inference runtimes, and low-latency systems engineering, you will thrive in this role.
Responsibilities
- Inference Pipeline & Architecture: Design, build, and maintain scalable, low-latency, and cost-effective model serving architectures for real-time and batch AI inference.
- Model Optimization & Compression: Implement model acceleration techniques including quantization (INT8/FP8/INT4, AWQ, GPTQ), pruning, distillation, and kernel optimizations.
- Serving Engines & Frameworks: Deploy and tune modern inference engines and runtimes (e.g., vLLM, TensorRT-LLM, ONNX Runtime, Triton Inference Server, TGI).
- Infrastructure & Compute Efficiency: Manage GPU memory utilization (KV caching, PagedAttention, batching strategies), optimize TTFT (Time to First Token) and token throughput, and profile hardware compute performance.
- System Integration & Telephony/APIs: Build high-concurrency API microservices (gRPC/REST/WebSockets) and integrate inference backends seamlessly with downstream business logic and media streaming pipelines.
- Benchmarking & Observability: Establish benchmarking frameworks for latency, throughput, cold starts, and accuracy degradation; implement robust monitoring and alerting for inference clusters.
Qualifications
Must-Have (Minimum Requirements):
- Education & Experience: Bachelor’s or Master’s degree in Computer Science, Electrical/Computer Engineering, or equivalent practical experience, with 3+ years of experience in AI/ML engineering, MLOps, or systems engineering.
- Programming Skills: Strong proficiency in Python and C++ / CUDA (or systems-level performance tuning).
- Deep Learning Frameworks: Hands-on experience with PyTorch internals, TorchScript, or ONNX.
- Inference Serving Tools: Proven experience deploying models with at least two of the following: vLLM, Triton Inference Server, TensorRT / TensorRT-LLM, DeepSpeed-FastGen, or TGI.
- Cloud & Containerization: Proficiency with Docker, Kubernetes, and orchestration of GPU-accelerated workloads (NVIDIA GPUs, CUDA drivers, NCCL).
- Performance Profiling: Familiarity with profiling tools (e.g., NVIDIA Nsight, PyTorch Profiler) to identify compute and memory bottlenecks.
Nice-to-Have (Preferred Qualifications):
- Experience with real-time streaming audio/voice AI pipelines (ASR, TTS, Voice Agent turn-taking) or Computer Vision at the edge.
- Experience in custom CUDA kernel development or Triton DSL.
- Strong understanding of distributed inference strategies (Tensor Parallelism, Pipeline Parallelism) across multi-GPU / multi-node setups.
- Experience in cloud cost optimization for large-scale GPU clusters.