DT
NLP & Machine Learning Engineer
Datago Technology LimitedPosted 7 hours ago
Is this job info correct?
Role
Datago turns Chinese and English financial text into structured data products: sentiment scores, named-entity recognition, event classification, used by institutional clients. The team fine-tunes BERT-family encoders on proprietary corpora, maintains rigorous evaluation standards, and builds pipelines around them. As the models are delivered to paying clients, evaluation standards are strict and all results must be auditable.
This is a model development role rather than a prompt-engineering position. The work spans training loops, dataset preparation and evaluation design, and the engineer is accountable for the figures the models produce. AI assistants are in the workflow but responsibility for results remains with the engineer.
Responsibilities
Benefits
Datago turns Chinese and English financial text into structured data products: sentiment scores, named-entity recognition, event classification, used by institutional clients. The team fine-tunes BERT-family encoders on proprietary corpora, maintains rigorous evaluation standards, and builds pipelines around them. As the models are delivered to paying clients, evaluation standards are strict and all results must be auditable.
This is a model development role rather than a prompt-engineering position. The work spans training loops, dataset preparation and evaluation design, and the engineer is accountable for the figures the models produce. AI assistants are in the workflow but responsibility for results remains with the engineer.
Responsibilities
- Fine-tune and maintain encoder models (BERT, and more) for classification, NER and sentiment tasks.
- Design evaluation frameworks that reflect client priorities: ranking consistency, correlation against ground truth, and strict point-in-time splits.
- Manage data annotation workflows end to end, including guidelines and quality control.
- Curate large text corpora, covering sampling, cleaning, deduplication and versioning.
- Optimize model serving where required, via ONNX export, quantization and batch inference.
- Maintain experiment records in W&B or an equivalent tool, logging each run's configuration, metrics and artifacts so results can be reproduced and compared.
- Maintain bleeding-edge knowledge in relating areas.
- 1-4 years of ML/NLP experience, or a research master's degree with substantive projects, coursework certificates alone are not considered sufficient.
- Hands-on Python and PyTorch proficiency: the candidate should have personally fine-tuned a transformer model and be able to explain the problems encountered and how they were identified.
- The HuggingFace transformers ecosystem, plus fluent pandas/numpy.
- Proficiency with the Linux command line and Git, including running experiments on remote GPUs.
- A critical approach to metrics: an understanding of why accuracy alone can mislead, what data leakage looks like, and why the evaluation split matters.
- Fluent Cantonese, working proficiency in English for documentation and code, ability to process Chinese text in both Traditional and Simplified forms.
- A degree in Computer Science, Mathematics, Statistics, Linguistics or a related discipline, or equivalent demonstrable experience.
- Finance or fintech domain exposure.
- LLM fine-tuning (LoRA/QLoRA), prompt evaluation, harness-use.
- SQL/ClickHouse.
Benefits
- Competitive salary
- Annual leave and group medical insurance
- Good team culture
- Assistance to apply for an IANG VISA