Junior Data Engineer
- Hiring from
- France
- Work type
- Hybrid
- Posted
Show job descriptionHide job description
About GitGuardian
GitGuardian is a global cybersecurity scale-up. The company is based in Paris, New-York City, Boston.
Among our early investors who saw our market value proposition, are the co-founder of GitHub, Scott Chacon, along with Solomon Hykes, Docker's co-founder. American and European top-tier VC firms have also invested in GitGuardian who raised its series C early 2026.
GitGuardian leads the way in securing the credential layer, helping organizations protect the secrets that let code, machines, and AI agents access systems with broad detection, identity context, and remediation at scale. Already trusted by 600K+ developers worldwide and by 50 of the fortune 500!
About your team and your mission
You will join the Data & AI Engineering team, whose mission is to build the data foundation and the internal AI-engineering capability that let GitGuardian run as an AI-native company.
The team works across three pillars. On the data side, we run the data platform, the business data models, and the ingestion chain behind every team's reporting and our product analytics, with the aim of data that is used, trusted, and increasingly near real-time. On internal AI tooling, we build the shared skills, internal MCPs, and the Analytics Agent, along with the governance and evaluation standards that make these tools safe to build on. And on apps, we maintain the internal app framework that lets anyone in the company ship an application to software engineering standards, with SSO, CI/CD, and managed hosting built in.
Data is already the backbone of the company's reporting and decision process. AI tooling and internal apps are newer ground, with real freedom to shape how all business stakeholders use data on the day-to-day.
Key challenges :
Build the data foundations for AI agents. AI agents are becoming the main way people query data. You'll help design the next generation of our warehouse and data models so that agents can understand them, trust them and answer questions correctly: clear semantics, well-documented metrics and structures built for machines as much as for humans.
Expand our data coverage to sensitive domains. Finance and HR data are joining our scope. You'll help bring them into the warehouse with the right level of rigor, access control and confidentiality.
Strengthen the existing data domains. Marketing, Sales and Product Analytics already rely on us. We want to make their data more complete, reliable and easy to use, on Snowflake today and ClickHouse soon.
Make data part of everyone's daily work. Our goal is for every team at GitGuardian to use Data & AI tools every day. That depends on trustworthy data models and clearly defined business metrics.
Keep the platform running smoothly while it grows. Day-to-day data requests and run topics matter. Handling them well lets the whole squad keep delivering on the larger projects of a 12–18 month roadmap.
Your responsibilities :
Write data models as code daily (SQL, Python with Snowpark and PySpark) to deliver new data for business insights.
Own your topics end to end, from ingestion to consumption by business users.
Work with internal stakeholders (Sales, Marketing, Finance, HR, Product) to understand their needs and turn them into well-defined business metrics.
Build and orchestrate pipelines with Dagster, and deploy them with Docker and Terraform on AWS.
Take part in the run: monitor pipelines, investigate data issues, and answer data requests from across the company.
Keep the code base healthy through code reviews, tests and documentation.
Work closely with the squad's AI Engineers so that our data powers the company's internal AI tools.
Technical environment
Coding Languages: Python (Snowpark & PySpark), SQL
Analytics databases: Snowflake, ClickHouse
Visualisation: Metabase
Orchestration: Dagster
Deployment: Docker, Terraform
Cloud Provider: AWS
VCS: GitLab
CI/CD: ArgoCD
About you
If you think you match at least 70% of these criteria, please apply!
Here's what we consider essential for success in this role:
A first experience in data engineering, such as an internship, apprenticeship or up to about 2 years in a role.
Strong SQL and good Python skills.
An understanding of data modeling concepts (dimensional modeling, fact and dimension tables, business metrics).
The ability to find your way in an existing, complex code base and learn from it quickly.
Comfort talking with non-technical stakeholders, asking the right questions and explaining your choices.
A high standard for quality: you care about data that is correct, tested and documented.
Fluent English in an international environment.
Availability to work from our office 3 days a week.
The following skills would strengthen your application but aren't required:
Experience with Snowflake and/or Clickhouse
Familiarity with agentic systems or LLM-based architectures applied to data pipelines.
Experience in high-growth startups or scale-ups.
The interview process
At GitGuardian, we are committed to building a diverse, equitable and inclusive workforce.
We will ask for your gender identity on the application page to help us understand the diversity of our applicant pool and to track our progress in attracting and hiring a diverse workforce. The information is optional and will not be disclosed to the hiring manager or the interview team and will not be considered in the hiring process. We appreciate your willingness to share this with us so that we can continue to improve our diversity, equity and inclusion efforts.
1. Manager interview
To discover your professional project and evaluate if there could be a mutual match
2. Technical interview with the team (1h30)
Purpose: We validate your hard skills and give you the opportunity to discuss with other staff engineers and data engineers
Skills Assessed: Python coding proficiency, Data architecture skills and overall communication and reasoning.
3. Engineering leadership interview
Purpose: To know more about yourself and your achievements, and present to you the team.
Skills Assessed: We assess your soft skills (ownership, communication) and motivation.
4.1 Final interview with Eric our CEO, onsite
Purpose: To detail our company’s vision and ambitions for the next couple of years.
You will also meet some team members for a coffee.
4.2 References check
You can start thinking about two contacts who can attest to your previous or current professional experiences. These contacts should be as recent as possible, and we will call them at the end of the process.