NLP / LLM Engineer — Mid-Senior | Full-time | Onsite | Visa Sponsorship Available
ABOUT HUMBLEBEEAI
We specialize in Computer Vision, Generative AI, and Business Analytics. An AI studio with tools, talent, and purpose. We build from our own AI ecosystem to deliver scalable, ethical, and lasting solutions for clients worldwide.
Our offices are located in Incheon (Bee Intelligence Global Co., Ltd, Republic of Korea), and in Tashkent and Namangan (Republic of Uzbekistan).
ABOUT THE ROLE
We are looking for an NLP/LLM Engineer to build and improve production-grade AI systems across RAG, agents, evaluation, fine-tuning, document processing, memory, speech, and observability.
You should be comfortable taking an LLM feature from experimentation and dataset preparation through evaluation, deployment, monitoring, and continuous improvement.
TEAM AND REPORTING
- You will report directly to the Chief Technology Officer.
- You will work day to day with a group of 4–5 engineers, owning the technical direction and delivery of the LLM workstream.
- The wider company team is 20–25 people across the Korea and Uzbekistan offices.
WHAT YOU WILL DO
- Design and improve production RAG pipelines, retrieval, chunking, embeddings, reranking, and context construction.
- Build multi-step LLM workflows and agents using frameworks such as LangChain, LangGraph, CrewAI, LlamaIndex, or similar.
- Develop query rewriting, intent classification, routing, structured extraction, tool calling, and context-sufficiency checks.
- Improve semantic, keyword, hybrid, filtered, and graph-based retrieval.
- Reduce hallucinations and improve groundedness, factual consistency, and answer quality.
- Build evaluation datasets and automated evaluation pipelines for retrieval and generation.
- Measure retrieval with Recall@K, Precision@K, MRR, and nDCG, and evaluate generation for correctness, relevance, groundedness, safety, latency, and cost.
- Fine-tune and adapt models using SFT, LoRA, QLoRA, PEFT, preference optimization, and distillation where useful.
- Work with OCR, PDF/DOCX processing, speech-to-text, conversational memory, and multilingual NLP.
- Trace and monitor model calls, retrieval, token usage, latency, failures, and cost.
- Run controlled experiments, regression tests, A/B tests, and human evaluations.
- Optimize systems for reliability, concurrency, streaming, caching, retries, and provider failures.
REQUIREMENTS — MUST HAVE
- 3+ years of experience in NLP, ML, Information Retrieval, or Applied AI.
- Strong Python engineering skills and experience building production services.
- Hands-on experience shipping LLM, RAG, or agent systems to production.
- Strong understanding of transformers, tokenization, embeddings, context windows, prompting, structured outputs, and decoding.
- Strong knowledge of semantic search, hybrid retrieval, metadata filtering, reranking, chunking, and indexing.
- Experience with LangChain, LangGraph, CrewAI, LlamaIndex, or comparable orchestration frameworks.
- Experience with vector or search systems such as Milvus, Pinecone, Qdrant, Weaviate, or Elasticsearch.
- Experience with OpenAI, Gemini, Hugging Face Transformers, and PyTorch.
- Practical knowledge of SFT, LoRA, QLoRA, and PEFT.
- Experience designing automated and human-in-the-loop LLM evaluations.
- Strong understanding of hallucination mitigation, prompt injection, PII and data leakage, and adversarial testing.
- Experience with FastAPI, REST APIs, Docker, Git, testing, and CI/CD.
- Experience with multilingual NLP systems and low-resource language challenges.
- Professional English proficiency.
NICE TO HAVE
- Experience training embedding or reranking models.
- Knowledge of BM25, cross-encoders, bi-encoders, and Reciprocal Rank Fusion.
- Experience with DPO or other preference-alignment methods.
- Experience serving models with vLLM, TGI, Triton, or similar.
- Experience with quantization, batching, KV caching, and GPU inference optimization.
- Familiarity with RAGAS, DeepEval, promptfoo, LangSmith, Langfuse, or custom evaluation systems.
- Experience with OCR and STT evaluation using WER/CER.
- Experience with Graph RAG, knowledge graphs, or synthetic dataset generation.
- Research publications, open-source contributions, or relevant competition experience.
- Experience mentoring or technically guiding a small engineering team.
WHAT WE OFFER
- End-to-end ownership of real LLM systems: RAG, agents, evaluation, fine-tuning, and deployment.
- Direct collaboration with the CTO, founders, and product leadership.
- High-performance compute, tooling, and infrastructure support.
- Compensation based on experience, visa sponsorship support, and a long-term growth trajectory.
INTERVIEW PROCESS
1. Document screening — review of your CV, portfolio, and application form.
2. Online technical interview — Zoom session covering your NLP/LLM experience and system design.
3. In-person interview — final conversation at one of our offices.
HOW TO APPLY
Apply here: https://docs.google.com/forms/d/e/1FAIpQLSen1LYdeGS_hr2EMnKB7IKuZDwd8pueVz8gdjs_nUfB1HSTUA/viewform
Questions: [email protected]
Please include: your CV or LinkedIn profile, relevant GitHub repositories/projects/publications, and a short description of one LLM or NLP system you shipped to production, including how you evaluated and improved it.
HOW TO WRITE YOUR CV
• Focus on the projects you have actually worked on, and describe your specific roles and contributions in detail.
• Rather than simply listing project names or results, try to convey the overall context of your experience by referring to the points below:
• How you identified and resolved issues that arose during system operation or development
• If you have experience applying your work to a real service and achieving measurable improvements, please include the results in numerical form (if disclosure is sensitive, you may omit that part).
• If you have personal projects or portfolios (e.g., GitHub), please include them as well. They will help us better understand your practical experience and capabilities.
HumbleBeeAI is an equal opportunity employer. We consider all qualified applicants without regard to age, gender, nationality, ethnicity, religion, disability, marital status, or any other characteristic protected by applicable law.
Applied Scientist II - Moloco Commerce Media
Seoulstart
Chief Operating Officer
HumbleBeeAI
PTW (Permit to Work) Coordinator
Bv
Renewables Permitting Team_Solar Energy
Solutions for Our Climate (SFOC)
[KOTRA Silicon Valley] Global Talent Pool
KOTRA Global Talent Center
Compute Country Lead, Korea
Anthropic