Location: Ghent, Belgium. ONTOFORCE helps life sciences organizations accelerate research and drug development by unlocking hidden insights from complex data. Our flagship platform, DISQOVER, is a life sciences data and knowledge platform built on semantic technology and knowledge graph foundations, designed to connect siloed internal, licensed, and public data in one searchable environment. Why this role exists DISQOVER is only as good as the data beneath it. A large part of that data comes from the public life sciences landscape: clinical trial registries, chemistry and target databases, scientific literature, regulatory sources, ontologies, and terminologies. These are dozens of sources, each with its own release cadence and schema conventions, and each capable of changing without notice. Ingesting these sources is the straightforward part. The greater challenge is making them agree with one another. The same gene, drug, disease, or organization may be named several different ways across sources, yet customers expect to search once and find all of it. Mapping those references, harmonizing them, and correcting the semantic meaning behind them is central to what DISQOVER does. It is what separates a search engine over documents from a genuine discovery platform. This role owns the machinery that produces this, not simply keeping it running, but making it predictable, observable, and trustworthy enough that customers never have to think about it. What you'll own The public data pipelines, end to end. You will be accountable for our public feeds landing correctly and on time, release after release. That includes the less visible work: handling failures, managing reprocessing, and responding when a source changes its export format without warning. It also brings you closer to the harmonization than it might sound. A great deal of the mapping and correction logic lives inside these pipelines, so improving how a source is ingested directly improves how well DISQOVER resolves meaning across all of them. Keeping the lights on and moving the needle are the same work here, not competing priorities. An early-warning system for upstream change. Today we detect these changes manually. You will help us find better and faster approaches: schema drift detection, statistical checks on volumes and distributions, diffing across releases, and monitoring source announcements. A change upstream should surface as a signal in our systems, not as a question from a customer. Quality as an engineering practice. Our data work does not yet have the testing and validation discipline we apply elsewhere in engineering, and we want to change that. You will define what "good" means for a feed, make it measurable, and ensure it is enforced automatically. You choose the approach and see it through, and what you decide becomes how the team works. Bringing the outside in. We run most of our transformation logic on an ETL engine we built in house, and the reason is provenance. It tracks where every value came from at column and cell level, which the mainstream stack does not offer without significant effort, and which matters a great deal to customers who need to trace a conclusion back to its source. That requirement is why we built rather than bought, and it is why the decision has held. It also means we cannot simply adopt tooling that treats lineage as an afterthought. We are looking for someone who brings real experience of the modern technical data stack and can tell us where the ecosystem has caught up with us, where it has not, and which ideas are worth borrowing even when the tools themselves are not. We value a clear argument and a migration path over a preference. That perspective is also market intelligence: our customers' data teams work with these tools, and their expectations should help shape where DISQOVER integrates. How we work ONTOFORCE is small enough that individual contributions are visible and genuinely matter. With that comes real ownership, alongside the expectation that you pick up what needs picking up and help a colleague past a blocker. Teamwork is not a slogan here; it is how the work gets done. This role suits someone who wants a say in how the team operates, rather than a tightly scoped remit with little contact beyond it. You will join a team that already runs this stack. You are not inheriting an abandoned system, and you will not be the only person who understands it. We use AI coding assistants daily and expect you to as well. What interests us is not enthusiasm but judgement: knowing where they accelerate the work, where they produce confident but incorrect output, and how you tell the difference. Seven or more years building and operating production data pipelines, including time as the accountable owner of a data domain rather than a contributor to someone else's, with a proven track record of improving performance and maintainability. Hands-on experience with the modern technical data stack. The specific tools matter less than the perspective: orchestration (Dagster, Airflow, Prefect, or Temporal), transformation and testing tooling, and lineage and cataloguing. We deliberately do not run most of this in house, which is precisely why we want someone who has. Strong software engineering habits, not just scripting. You structure code that others can maintain, you test it, and you understand what makes performant code performant. We write a lot of Python, so comfort with it matters, though we care more about how you build than which language you used before. Much of our transformation logic runs on our own provenance-tracking engine rather than in SQL, so SQL fluency is welcome but not essential. Demonstrated experience implementing automated data quality checks and testing frameworks. Comfort working with Kubernetes. Our public data pipelines run entirely on EKS, so editing deployments, reading logs, and reasoning about Jobs and resource limits will be part of your routine. You do not need to know how to stand up a cluster, but you should be at home operating inside one. The instinct to ask "how would I know if this were wrong?" and a practice you introduced as a result that outlasted your involvement, whether in testing, review, monitoring, or documentation. To us, "senior" means the improvements you make remain in place after you move on. An architecture or build-versus-buy decision you owned, and enough time living with the consequences to judge honestly whether you got it right. Clear communication with non-engineers, the confidence to disagree with us when you have evidence, and openness to customer input. A deeply technical, analytical enthusiasm for the insights data can deliver. A process-oriented mindset, with the understanding that scaling depends on continuously improving how the work is done. Genuine curiosity about the wider data engineering landscape. You keep an eye on emerging tools, frameworks, and approaches, and you are eager to try them, not for their own sake, but to see whether they genuinely improve how you work. You bring a critical, impact-driven mindset to new technology. The following would strengthen an application, though none are required: Some exposure to applying ML or LLMs to matching, classification, or entity-resolution problems. Curiosity and familiarity are enough; we are not looking for a research background. Life sciences or biomedical data experience. Knowledge graphs in either tradition: RDF and SPARQL, or labelled property graphs. Experience with ontologies and controlled vocabularies. Semantic layer or master data management work in any domain. Data contracts, or vector search and embedding-based retrieval. We have yet to meet someone strong across all of this, and we do not expect to. Tell us which parts you are strong in and which you would like to grow into.
Data Engineer
IVC Evidensia
Tender Specialist and Business Support – Dutch & French Speaker
Agilent
Senior Account Executive, Customer Base
Workday
Presales Solution Architect – Defense Center
Orange Cyberdefense
SSA Senior Engineer
Westinghousenuclear
Cost Manager
AECOM