NLP Data Scientist
FyerxJob Description
This is a remote position.
NLP Data Scientist / LLM Fine-Tuning Specialist
Job Details
- Employment Type: Contract
- Work Mode: Remote
- Location: Offshore
- Total Experience Required: 4 to 8 years
- Relevant Experience Required: 3+ years of dedicated natural language processing (NLP) and hands-on Large Language Model (LLM) fine-tuning experience
- Mandatory Certification: Google Cloud Certified Professional Machine Learning Engineer or AWS Certified Machine Learning - Specialty
Job Summary
We are seeking an experienced NLP Data Scientist / LLM Fine-Tuning Specialist to take ownership of our specialized open-source model optimization tracks. The ideal candidate will possess deep expertise in deep learning, dataset preparation, and parameter-efficient training methodologies to fine-tune foundational models for industry-specific terminology, domain-specific reasoning, and custom task execution.
Key Responsibilities
- Lead LLM fine-tuning initiatives, leveraging Parameter-Efficient Fine-Tuning techniques (PEFT) including LoRA, QLoRA, Prefix Tuning, and Prompt Tuning to optimize open-source architectures (e.g., Llama, Mistral).
- Curate, clean, and structure high-quality training datasets, implementing automated data deduplication, tokenization schemes, synthetic data generation pipelines, and human-in-the-loop validation frameworks.
- Implement advanced reinforcement learning alignment layers, configuring Reinforcement Learning from Human Feedback (RLHF) or Direct Preference Optimization (DPO) to enforce model safety, helpfulness, and tone guardrails.
- Optimize model footprint constraints and memory overhead, applying post-training quantization techniques (e.g., GGUF, AWQ, GPTQ) to minimize parameter degradation and compute budgets.
- Design rigorous evaluation benchmarks and metrics panels, executing automated validation tests (e.g., BLEU, ROUGE, custom verification matrices) to audit model hallucinations, factual accuracy, and domain alignment.
- Manage distributed deep learning training jobs, scaling pipeline configurations, tensor parallelism parameters, and gradient checkpointing scripts across multi-GPU compute blocks.
- Collaborate with MLOps infrastructure teams, formatting completed model weight checkpoints cleanly for scalable cloud deployment and real-time inference serving layers.
Requirements
- 4 to 8 years of core data science or advanced machine learning engineering experience, with 3+ dedicated years actively training, evaluation, and fine-tuning natural language processing systems.
- Expert-level technical mastery of Python, deep learning frameworks (PyTorch), transformer architectures (Hugging Face Transformers, Accelerate, PEFT), and vector calculations.
- Deep structural understanding of attention mechanisms, tokenization constraints, context window degradation behaviors, loss function optimization, and hardware compute limitations (CUDA).
- Mandatory certification: Professional ML Engineer or Specialty Machine Learning credential from a major cloud vendor (AWS/GCP).
Preferred Qualifications
- Master’s or Ph.D. in Computer Science, Data Science, Computational Linguistics, or an adjacent quantitative field with a research focus on neural network text models.
- Prior experience implementing custom embedding model structures or optimizing domain-specific classification layers inside constrained enterprise runtimes.