岗位描述
Pillar 1: Model Hosting & Inference Optimization Design, build, and maintain scalable, reliable model hosting systems to support large-scale deployment of AI models (LLMs, embedding models, STT/TTS, etc.) across diverse hardware platforms. Develop and implement inference optimization strategies to achieve low latency, high throughput, and cost-efficiency, including model quantization, KV cache optimization, and dynamic batch processing / Continuous Batching. Evaluate, integrate, and customize state-of-the-art inference frameworks (vLLM, TensorRT-LLM, SGLang, etc.) to optimize model performance on target hardware. Monitor and maintain inference system health, track key performance metrics (latency, throughput, TTFT, memory usage, availability), and troubleshoot performance bottlenecks or deployment issues. Collaborate with hardware teams to leverage hardware-specific and optimize hosting infrastructure for maximum resource utilization. Ensure model hosting systems comply with reliability, scalability, and security standards, supporting high-availability production workloads. Design and build end-to-end, scalable model fine-tuning pipelines to support domain-specific adaptation of pre-trained using custom, domain-specified datasets. Collaborate with data scientists and domain experts to understand domain requirements, define fine-tuning objectives, and validate fine-tuned model performance against domain-specific metrics. Bachelor's/Master's/PhD degree in Machine Learning, Natural Language Processing, Computer Science, Data Science, Statistics, or related fields. 3+ years' experience leading teams, including people management and delivery ownership. 5+ years of experience in AI, with proven expertise in both model inference optimization/hosting and model fine-tuning pipeline development; experience with LLM-related work is a strong plus. Proficiency in programming languages such as Python, CUDA; familiarity with GPU/CPU architecture and high-performance computing (HPC) principles. Deep understanding of AI model inference: KV cache management, batch processing, quantization (INT4/FP8/GPTQ/AWQ), operator optimization, and inference framework integration (vLLM, TensorRT-LLM, SGLang, etc.). Hands-on experience building and maintaining model hosting systems, including knowledge of containerization (Docker, Kubernetes) and cloud computing platforms (AWS, GCP, Azure) for scalable deployment. Expertise in designing end-to-end model fine-tuning pipelines: data preprocessing, distributed training, hyperparameter tuning, and integration with pre-trained models Familiarity with fine-tuning frameworks and tools (Hugging Face Transformers, Accelerate, LoRA, QLoRA) and experience optimizing fine-tuning workflows for efficiency. Experience with performance benchmarking, monitoring, and troubleshooting for both inference and fine-tuning systems. Demonstrates an AI-native mindset, applying AI-driven approaches to improve productivity, quality, and decision-making. Experienced in leveraging coding assistants (e.g., AI pair-programming tools) to accelerate software development, enhance code quality, and support engineering best practices.