Junior Data Engineer - Jordan Office
Job Overview
As a Junior Data Engineer at Ligadata, you will support the development, maintenance, and monitoring of data pipelines and Big Data solutions. You will work with senior engineers and cross-functional teams to ensure reliable data processing, data quality, and timely delivery.
The role requires a good foundation in SQL, Linux, Shell scripting, data analysis, and Big Data technologies, with a strong willingness to learn and troubleshoot within a production data environment.
Responsibilities
- Develop and maintain ETL/ELT data pipelines.
- Write and optimize SQL queries for data processing, validation, and analysis.
- Support data workflows using Apache Airflow.
- Work with Big Data technologies such as Hadoop, HDFS, Hive, Spark, Presto/Trino, Kafka, and HBase.
- Perform data validation, reconciliation, and data-quality checks.
- Develop scripts and automation using Shell/Bash and Python.
- Monitor data pipelines and assist in troubleshooting job failures and production issues.
- Work with structured and semi-structured data formats such as Parquet, JSON, CSV, and Avro.
- Support applications and data workloads running on Kubernetes (K8s).
- Analyze data to identify inconsistencies, anomalies, and operational issues.
- Participate in code reviews, documentation, and continuous improvement activities.
- Collaborate with senior engineers, DevOps, QA, database, and analytics teams.
- Use AI-assisted engineering tools to support development, troubleshooting, documentation, and data analysis while validating generated results before use.
Qualifications
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related field.
- 1–3 years of experience in Data Engineering, Big Data, Software Engineering, or a related role.
- Good knowledge of SQL and relational database concepts.
- Good understanding of Linux and Shell/Bash scripting.
- Basic knowledge of the Hadoop ecosystem, including HDFS and Hive.
- Familiarity with Spark, Presto/Trino, and Apache Airflow.
- Basic understanding of Kafka and distributed data-processing concepts.
- Familiarity with Kubernetes and containerized environments.
- Knowledge of at least one programming language, preferably Python, Scala, or Java.
- Basic understanding of ETL/ELT, data warehousing, data quality, and data analysis.
- Familiarity with Git and software-development practices.
- Knowledge and practical experience using AI tools such as ChatGPT, GitHub Copilot, or similar tools for engineering tasks.
- Good analytical, troubleshooting, and problem-solving skills.
- Strong willingness to learn and develop technical skills.
- Good communication and teamwork skills.
- Self-motivated with a strong sense of ownership.