Client: Agency of the United Nation
Location: Italy / Remote
Duration: Initial 3 months, SEP – NOV 2026 with possible extension subject to continuing business need, available budget and vendor performance
Duration: Based on Experience and budget availability
Organizational context
The client is a large international organization seeking to deliver a well-engineered, AI-augmented enterprise search capability to improve access to knowledge across its systems and support its global operations.
Duties and responsibilities
Under the supervision of the ICT Solutions Coordinator, the incumbent will work closely with ICT teams to deliver a well-engineered functioning enterprise search with Artificial Intelligence (AI) augmentation based on Enterprise Elastic search.
Duties include:
Data Engineering:
· Design and implement scalable data ingestion pipelines and connectors to ingest structured, semi-structured, and unstructured content from enterprise sources (SharePoint, Liferay, Web crawls, Data Lake, Corporate Systems, etc.) into Elasticsearch or an equivalent search index, supporting batch, incremental, and near-real-time indexing processes.
· Design mechanisms to track document versions, source provenance, access permissions, timestamp updates, and deletion events to keep the search index accurate and current; develop content extraction pipelines for PDF, Word, Excel, PowerPoint, HTML, emails, scanned documents, and other enterprise formats, converting them into standard markdown and/or vector embeddings to improve AI readability.
· Design and implement semantic chunking strategies and hybrid search logic optimized for retrieval quality (chunk size, overlap, section-aware splitting, heading preservation, table handling, context retention), and implement metadata extraction, enrichment, and deduplication of content during ingestion.
2. Build Retrieval Capabilities:
· Develop hybrid search capabilities combining keyword-based search, semantic vector search, metadata filtering, and contextual retrieval, complemented by re-ranking pipelines using specialized embedding models, ranking logic, or other suitable re-ranking techniques to improve relevance of retrieved results.
· Implement advanced retrieval techniques such as query rewriting, query expansion, function or tool calling, multi-query retrieval, metadata-aware retrieval, parent-child retrieval, contextual document embeddings, and contextual compression, along with security controls to ensure users can only retrieve and access relevant information they are authorized to view.
3. Build RAG Pipelines:
· Design and build the Retrieval-Augmented Generation (RAG) pipeline that retrieves relevant enterprise content and uses large language models and/or specialized AI models to generate grounded answers, including agentic workflows where the AI application can invoke tools, perform multi-step reasoning, call enterprise APIs, refine searches, and retrieve additional context to answer user queries.
· Implement prompt engineering and orchestration patterns for reliable and relevant response generation, including system prompts, retrieval prompts, guardrails, context assembly, and response formatting, along with fallback strategies for insufficient context, ambiguous questions, or low-confidence retrieval results.
Professional requirements
· Elasticsearch engineering. Proven and deep hands-on experience with query DSL, BM25 tuning, function_score, boosting/decay functions, and multi-field matching strategies amongst other Elasticsearch features.
· Index & data modelling architecture. Ability to design mappings, choose the right field types, and configure custom analyzers/tokenizers per content type (e.g., code vs. prose vs. structured records vs. multimedia content).
· Connector / ingestion pipeline. Real experience building or configuring pipelines for SharePoint, Liferay, databases, and Azure Data Lake — including incremental sync/CDC, handling deletes/updates, and dealing with rate limits and API intricacies of each source.
· Hybrid & semantic search (lexical, vector, ELSER or similar).
Proven experience of implementing hybrid search including a deep understanding of semantic search and keyword search.
· Performance, scaling & cluster operations. Shard strategy, index sizing, reindexing strategy, query latency tuning, and general cluster health management
· Search evaluation & relevance testing methodology. Experience of building a ground truth and gold-standard benchmark of relevant samples and test queries with expected results, measure precision/recall and/or NDCG and similar evaluation metrics, and iterate against it.
· Proficiency in Python and experience with data processing frameworks and libraries.
· Strong hands-on experience with Elasticsearch, OpenSearch, Azure AI Search, or similar enterprise search platforms.
· Experience implementing semantic chunking strategies that split content into contextually coherent sections while preserving headings, structure, metadata, and parent-document relationships to improve retrieval accuracy in RAG applications.
· Experience designing ingestion pipelines that convert enterprise documents into structured Markdown using tools such as Marker, Docling, or equivalent document-conversion frameworks to preserve layout, tables, headings, and metadata for downstream RAG indexing.
· Experience with embedding models, re-ranking models, cross-encoders, prompt engineering, context window management, and response grounding techniques.
· Experience with LLM orchestration frameworks such as LangChain, LlamaIndex, Haystack, or equivalent frameworks.
· Experience with tool calling, agentic workflows, function calling, multi-step retrieval, and AI application orchestration.
· Experience working with commercial or open-source LLMs, such as Azure OpenAI, OpenAI, Anthropic, Google Gemini, Meta Llama, Mistral, Jina, or similar models.
· Knowledge of React and/or similar front-end technologies and frameworks for web application development.
Qualification and experience
First level University degree IN Computer Science, Computer Engineering, Information Systems or related with at least 5 years of professional work experience, of which minimum 3 years hands-on experience building enterprise search, AI-powered search, semantic search, Retrieval-Augmented Generation, or LLM-based applications is required.
Languages
Excellent written and verbal communication skills in English essential.
Research Engineer
Ecosmic
Search Backend Engineer (M/F/D)
Witglobal
Search Platform Engineer (M/F/D)
Witglobal
Python Engineer (Remote)
Jobs Ai
Full-Stack Developer (Remote)
Jobs Ai
Senior Sharepoint Developer
Proxify