Talent.com
Phincon
Inference EngineerPhincon • Kota Administrasi Jakarta Selatan, DKI Jakarta, ID
Cari pekerjaan lain
Inference Engineer

Inference Engineer

Phincon • Kota Administrasi Jakarta Selatan, DKI Jakarta, ID
29 hari yang lalu
Uraian Tugas

We are looking for a skilled and passionate Inference Engineer to design, optimize, and deploy high-performance AI inference pipelines for production environments. In this role, you will bridge the gap between machine learning models and high-throughput, low-latency production infrastructure. You will be responsible for optimizing state-of-the-art models (such as LLMs, Speech/Audio models, and Computer Vision architectures), reducing compute costs, and ensuring resilient, scalable serving on enterprise-grade hardware (GPUs/accelerators). If you have deep expertise in model quantization, inference runtimes, and low-latency systems engineering, you will thrive in this role.

Responsibilities

  • Inference Pipeline & Architecture: Design, build, and maintain scalable, low-latency, and cost-effective model serving architectures for real-time and batch AI inference.
  • Model Optimization & Compression: Implement model acceleration techniques including quantization (INT8/FP8/INT4, AWQ, GPTQ), pruning, distillation, and kernel optimizations.
  • Serving Engines & Frameworks: Deploy and tune modern inference engines and runtimes (e.g., vLLM, TensorRT-LLM, ONNX Runtime, Triton Inference Server, TGI).
  • Infrastructure & Compute Efficiency: Manage GPU memory utilization (KV caching, PagedAttention, batching strategies), optimize TTFT (Time to First Token) and token throughput, and profile hardware compute performance.
  • System Integration & Telephony/APIs: Build high-concurrency API microservices (gRPC/REST/WebSockets) and integrate inference backends seamlessly with downstream business logic and media streaming pipelines.
  • Benchmarking & Observability: Establish benchmarking frameworks for latency, throughput, cold starts, and accuracy degradation; implement robust monitoring and alerting for inference clusters.

Qualifications

Must-Have (Minimum Requirements):

  • Education & Experience: Bachelor’s or Master’s degree in Computer Science, Electrical/Computer Engineering, or equivalent practical experience, with 3+ years of experience in AI/ML engineering, MLOps, or systems engineering.
  • Programming Skills: Strong proficiency in Python and C++ / CUDA (or systems-level performance tuning).
  • Deep Learning Frameworks: Hands-on experience with PyTorch internals, TorchScript, or ONNX.
  • Inference Serving Tools: Proven experience deploying models with at least two of the following: vLLM, Triton Inference Server, TensorRT / TensorRT-LLM, DeepSpeed-FastGen, or TGI.
  • Cloud & Containerization: Proficiency with Docker, Kubernetes, and orchestration of GPU-accelerated workloads (NVIDIA GPUs, CUDA drivers, NCCL).
  • Performance Profiling: Familiarity with profiling tools (e.g., NVIDIA Nsight, PyTorch Profiler) to identify compute and memory bottlenecks.

Nice-to-Have (Preferred Qualifications):

  • Experience with real-time streaming audio/voice AI pipelines (ASR, TTS, Voice Agent turn-taking) or Computer Vision at the edge.
  • Experience in custom CUDA kernel development or Triton DSL.
  • Strong understanding of distributed inference strategies (Tensor Parallelism, Pipeline Parallelism) across multi-GPU / multi-node setups.
  • Experience in cloud cost optimization for large-scale GPU clusters.

Buat peringatan pekerjaan untuk pencarian ini

Inference Engineer • Kota Administrasi Jakarta Selatan, DKI Jakarta, ID