About The Agency Fund Somewhere in East Africa right now, a field worker is getting coached by an AI trainer on how to run a behavior change program. In South Asia, a nurse is consulting a clinical copilot mid-shift. A smallholder farmer is asking a chatbot whether to plant this week. The Agency Fund (TAF) builds the AI systems behind those interactions — and then works to make sure they actually work for the people using them. We're a US non-profit that embeds engineers and product managers directly inside evidence-based nonprofits across East Africa, South Asia, and beyond, helping them build data infrastructure, run experiments, and scale impact through AI and product thinking. Our team of 22 combines expertise in software, AI, social psychology, and development economics. Our community of 200+ partner organizations reaches millions of people annually. We're entrepreneurial, flat, and operate with minimal hierarchy — everyone contributes to a shared mission in ways that are hard to replicate elsewhere. The Role We're looking for an AI/ML Engineer to join our team. You'll design, build, evaluate, deploy and improve AI/ML systems that power Agency Fund's platforms and partner applications, like evaluation infrastructure for non-profits, AI coaches for training field workers, copilots to support nurses and doctors, chatbots that provide advisory to farmers, and more. You will serve as an ML expert for the non-profits we fund and the partners we collaborate with, helping translate the best practices from AI research, especially in evaluating AI products, into production in the form of tools, playbooks (like https://eval.playbook.org.ai ) and frameworks. Your work will directly improve outcomes for millions of people served by our NGO partners. What You'll Do Contribute to tools supported by The Agency Fund like Calibrate , an open-source AI evaluation infrastructure for non-profits. Work closely with behavioral scientists and researchers to identify common problems and translate solutions/learnings into reusable tools, playbooks, frameworks and publications. Open-source any non-trivial innovations that come out of our in-kind work. Ship fast by implementing rapid development and deployment cycles to deliver solutions efficiently and iteratively. Document technical decisions and maintain engineering standards across AI components Stay current with applied AI research and bring relevant advances to the team. Provide advice and support to NGOs building generative AI products. Be a mentor to NGOs, demonstrating industry best practices, culture, and tooling to the staff we may work with, helping to grow these capabilities from within as well. Participate in organizing workshops and seminars on sharing the lessons, best practices and tools that emerge from our work. Conduct site visits to observe NGO operations and draft recommendations to address technical needs. Be agentic by fostering a culture of proactive problem-solving and initiative-taking. Who You Are Foundations Strong foundation in math, deep learning, LLMs and modern AI systems. Demonstrable proof of work: 4+ years of hands-on ML engineering experience where you have built, evaluated and systematically improved at least one end-user facing product powered by deep learning beyond integrating model APIs, or building internal dashboards or contributing to ML infra. At least 1–2 years of experience building LLM-powered applications, with at least one production-grade agent beyond proof-of-concept demonstrations. Hands-on experience evaluating AI systems with a deep understanding of dataset and evaluation design. Fluency with prompting and context engineering for agentic systems, and a clear sense of when the fix is a better prompt versus a better model, tool or pipeline. Must be very comfortable with programming in Python. This role is a very hands-on. Comfortable using coding agents to produce quality work, not AI slop. You take full accountability for what you ship with minimal need for additional verification. Comfort with reading research papers and quickly testing relevant ideas. How you operate Experimental mindset. You reason carefully about data, model and evaluation together, and make iterative, measurable progress rather than purely chasing hunches. Ability to think from first principles and design practical, scalable solutions. You care about making sure the product works for the intended users first. Any research artifact is a welcome side effect, not the goal. You identify what needs to be done and do it without waiting to be told. You take ownership and show urgency to see things through till the end Detail-oriented with a keen eye for spotting mistakes early. Strong written and verbal communication skills. Ability to communicate technical concepts clearly to non-engineering audiences Care deeply about building AI responsibly and with direct social benefit Bonus points if you have Experience with multilingual NLP or working with low-resource languages — this is where much of our partner work actually happens (Swahili, Hindi, and local languages across East Africa and South Asia), and candidates with this background will hit the ground running Prior work in global health, international development, or social impact technology Familiarity with behavioral science or designing AI for behavior change Experience with data annotation pipelines and evaluation methodology for LLMs Background in AI safety, alignment, or responsible AI frameworks Experience with cloud infrastructure (AWS, GCP, or Azure) and deploying models in production Publications at top AI conferences, or substantial technical writing about your work Why Join Us Compensation We're not competing with big-tech salaries and we'd rather be honest about that upfront; however, we are competitive within the mission-driven tech space. How we work Async-first, fully remote — work from wherever you are, on hours that fit your timezone 22-person team with minimal hierarchy and short feedback loops Work alongside behavioral scientists, development economists, field researchers, and embedded engineers — the ML problems here emerge from real operational context, not product roadmap speculation Open-source culture: your best work gets published, not locked inside a product What makes TAF's way of working unusual: Most ML roles sit you inside a product team with a defined roadmap. At TAF, the problems surface in the field — a nurse using a clinical copilot in Kenya, a field worker whose AI trainer is producing inconsistent outputs, a farmer who stops trusting the chatbot. You're close enough to those users to understand why, and empowered to change what you built as a result. There's no layer between you and the problem. The team is small by design. That means you own things end-to-end — from scoping the ML approach with a behavioral scientist, to deploying it, to sitting in on a site visit where you watch it fail in a way you didn't anticipate. That loop is what makes the work compound. Location & Hiring Remote (Sub-Saharan African or South Asia) Fully remote and globally distributed. We hire internationally on a contractor basis — there is no single office or headquarters timezone. We have team members across North America, Europe, East Africa, East Asia, and South Asia. Some timezone overlap with East Africa and South Asia is useful given where our partners are based, but we don't mandate core hours.
AI/ML Software Engineer-Task Creator (RL Environments) (Contract)
Careerflow
ML- AI Engineer
Thinkbridge
Data Security AI/ML Engineer
Nitka
Senior Machine Learning Engineer / Tech Lead - AI & ML
Civo
Python and Kubernetes Software Engineer - Data, Workflows, AI/ML & Analytics
Canonical
ML Engineer
npv labs