Skill diminta
AuditingFastAPINestJSNode.jsPostgreSQLPythonRedis
Deskripsi
We build production AI systems for enterprise clients across many industries in Indonesia, running our own GPU inference infrastructure end to end. We're hiring an
- AI Engineer
- to
- design and build AI architecture
- (RAG pipelines, agentic systems, and the serving stack behind them) and ship it to production.
- What You'll Do
- Build RAG pipelines end to end: ingestion, chunking, embedding, indexing, hybrid search (dense + sparse, RRF), reranking, and generation.
- Design agentic workflows with tool calling, routing, and state management, including connectors via MCP.
- Deploy and tune self-hosted LLMs (vLLM) on NVIDIA GPUs; optimize VRAM, batching, concurrency, and caching.
- Ship AI features into backend services (FastAPI, Celery, NestJS, Socket.IO / Redis) with ACL, audit logging, and evaluation.
- Work with PMs, backend, and frontend developers, and document in Bahasa Indonesia.
- Required
- Real LLM projects you can show and walk us through (work, internships, open-source, or personal).
- Strong Python (async); solid grasp of RAG, prompting, tool use, and agentic patterns.
- Hands-on with vector databases, embedding models, FastAPI, and Celery.
- Some exposure to GPUs and model serving.
- Nice to Have
- vLLM or Ollama and GPU optimization; Milvus and hybrid search / reranking.
- MCP, PydanticAI, LangChain, LangGraph; NestJS / Node, Socket.IO, Redis.
- Domain experience in enterprise knowledge platforms, medical/clinical AI, Human Capital / HR analytics, or AI CV / resume screening.
- Stack
PydanticAI · LangChain · LangGraph · FastAPI · Celery · Milvus · vLLM · Ollama · PostgreSQL · NestJS · Socket.IO · Redis