Lead the design, training, evaluation, and deployment of production-grade, on-premise AI systems with an emphasis on fine-tuned and multi-agent LLM solutions, safety/red-teaming, and scalable MLOps in secure or air-gapped environments. Work with open-source model families and local inference stacks to deliver reliable, secure, and cost-efficient services on-site in Bandung or/and Busan.
Design and implement multi-agent orchestration and tool-use pipelines (e.g., LangGraph/LangChain/AutoGen), including function calling, RAG, structured outputs, fallbacks, and recovery strategies.
Build rigorous red-teaming and safety evaluation harnesses; simulate jailbreaks, prompt injection, data exfiltration, and model manipulation; implement guardrails, policies, and moderation.
Conduct adversarial and robustness testing for NLP/CV models; assess distribution shift, perturbations, poisoning risks; implement mitigations and hardening.
Architect retrieval-augmented systems with vector databases; optimize chunking, embeddings, indexing, hybrid search, re-ranking, and latency for reliable grounding.
Own performance and cost optimization: quantization (GGUF, GPTQ, AWQ), batching, KV cache management, speculative decoding, caching, and GPU utilization.
Develop production APIs/services with FastAPI or gRPC; implement observability, tracing, canarying, and human-in-the-loop feedback loops; monitor quality drift and handle incidents.
Contribute to internal AI infrastructure, tooling, and reusable components; enforce reproducibility and governance with MLflow, model registries, and artifact stores.
Deploy and operate models on-prem (VMs/Kubernetes), including versioning, rollback, autoscaling, and secure upgrade paths for air-gapped sites.
Collaborate with product, engineering, and domain teams to scope experiments and deliverables; produce clear design docs, threat models, and runbooks.
Mentor junior engineers; drive best practices, code reviews, and knowledge sharing.
Demonstrated LLM fine-tuning experience: SFT, LoRA/QLoRA, DPO or RLHF; dataset preparation, synthetic data generation, and large-scale evaluation.
Self-hosted model experience with at least one open-source family (e.g., Llama, Qwen, Mistral) and on-prem inference stacks (vLLM, TGI, TensorRT-LLM, Ollama).
Security and safety practices: prompt-injection defenses, PII handling, RBAC, secrets management, audit logging; familiarity with regulated/on-prem environments and local data protection requirements.
Data lowongan bersumber dari glints. Tombol “Lamar” mengarahkan Anda ke halaman aslinya.