GES is seeking a contract data engineer to deliver the technical core of a commercial customer data validation project for a municipal water utility. The work is self-contained, file-based, and has a defined end point. GES retains project management and all client-facing responsibility; the successful candidate works to our engagement lead. The work The client maintains roughly 19,000 commercial customer accounts in a utility billing system. Those records were accumulated over many years across inconsistent business processes, and most lack reliable business classification codes and business contact information. The client cannot currently determine what kind of business sits behind each account, nor reliably reach that business during a service interruption. Your job is to fix that as a data exercise. You will take an extract of those accounts, match it against business license records and county parcel data, enrich the matched records with industry classifications and contact details, score the confidence of every match, identify duplicates, and produce a clean exception queue for the records that cannot be confidently resolved. No production system is modified. There is no integration build and no application development. The deliverable is files and documentation. What you will do • Profile and assess four source datasets for structure, completeness, and quality, and document what you find. • Standardize and normalize addresses across all sources. Service and mailing addresses are the primary matching key — usable coordinates are not available and one county field is unreliable. • Build a tiered matching approach — deterministic matching first, probabilistic linkage on the residual — and document the methodology so it can be reproduced by someone else. • Assign and calibrate confidence scores against thresholds agreed with the client. Records below threshold go to the exception queue with a stated reason; nothing is force-matched to hit a number. • Assign SIC and NAICS classifications where they can be established from approved sources, and identify duplicates and records requiring manual review. • Write the technical documentation — data dictionary, source-to-target mapping, transformation and business rules, and the methodology summary. A word on that last point, because it is the one candidates most often underestimate. The client's acceptance criteria weight documentation quality as heavily as match rate. The documentation must allow their staff or a future vendor to understand, validate, reproduce, maintain, and extend the work. In practice a substantial share of this engagement is written deliverables, not code. If writing is not something you enjoy, this is not the right role. Required • Production experience with record linkage or entity resolution. Not fuzzy string matching in a script — designing blocking strategies, choosing comparison levels, and calibrating a match model. Splink, dedupe, recordlinkage, or an equivalent you can defend. • US address standardization. libpostal, usaddress, USPS Publication 28 conventions, or comparable. Directional, suffix, and unit-designator handling. • Strong SQL and Python. DuckDB, polars, or pandas. This is moderate-scale work — tens of thousands of records against hundreds of thousands. Distributed computing experience is not required and not the point. • Data quality assessment. Profiling, duplicate analysis, and the ability to describe data problems clearly to a non-technical audience. • Technical writing. You will be asked for a sample — a data dictionary, mapping document, or methodology write-up you produced. • Explainable methods. Every match must be traceable to a documented rule or an interpretable score. Approaches that cannot justify an individual result are not acceptable for this client. Preferred • SIC and NAICS classification systems, including the 2022 NAICS revision and SIC crosswalks. • County parcel data and basic GIS handling — shapefiles, GeoJSON, spatial joins as flat reference data. • Utility customer information systems, billing data, or municipal government data work. • Microsoft Fabric or Azure data tooling. Architecture direction is being finalized; familiarity is an advantage, not a prerequisite. • Practical use of LLMs for classification or exception triage within a documented, reproducible pipeline. Eligibility — please read before applying These are client-mandated and not negotiable: • US citizenship and US-based work location. All data processing must occur within US-based infrastructure. Work performed outside the United States cannot be accepted under any circumstances. • E-Verify enrollment. If engaging as a company, your business must be enrolled in the federal E-Verify program and able to provide a notarized subcontractor affidavit with your E-Verify Company ID at contract execution, per Georgia O.C.G.A. § 13-10-91. • Confidentiality. A data confidentiality agreement is required. Client data may not be shared, retained, or used outside this engagement. Engagement terms Type - Contract / subcontract engagement Duration - Approximately 8 weeks from kickoff Effort - In the order of 100 hours across the engagement — roughly 12–15 hours per week, flexible Start - Early September 2026 Location Remote, within the United States. No travel anticipated. Rate - Please state your basis — hourly, daily, or fixed — against the effort above.
Sr. Data Engineer
Versa Networks
Senior Data Engineer
Chenega
Data Center Networking Engineer
Hpe
Sr AI Data Engineer
Geaerospace
Staff Software Engineer, Data Warehouse Foundation
Sr AI Data Engineer
Geaerospace