AI Infrastructure Engineer specializing in building and scaling the backend systems that power large-scale machine learning.
Bridging the gap between AI development and production — orchestrating distributed GPU clusters, tuning low-latency inference engines, and hardening core data infrastructure.
Skills
─4 competencies- 01Architecting high-performance GPU and AI clusters
- 02Optimizing serving engines for low-latency production inference
- 03Scaling distributed vector databases and data pipelines
- 04Automating reliable MLOps and infrastructure engineering
Stats
─Expertise
─6 offeringsGPU Cluster Orchestration
Provision and optimize multi-node GPU clusters using Kubernetes, managing resource allocation, VRAM utilization, and multi-tenant isolation.
High-Performance Inference Serving
Set up highly optimized serving engines like vLLM, TensorRT-LLM, and Triton to minimize TTFT and maximize token throughput for production workloads.
Vector DB & RAG Storage Systems
Deploy and scale distributed vector databases like Weaviate, Milvus, or Qdrant, optimizing indexing strategies and retrieval pipeline speeds.
AI Infrastructure Pipelines (MLOps)
Build robust CI/CD and data engineering pipelines, automating weight distribution, checkpointing, and dynamic cluster autoscaling.
Distributed Training Infra
Architect infrastructure setups for model fine-tuning and training, configuring data-parallel and model-parallel setups with Ray and DeepSpeed.
Compute & Cost Monitoring
Implement full-stack observability frameworks to track GPU metrics, prompt cache hit-rates, latency, and cloud compute expenditures.