We are looking for a passionate and detail-oriented Data Engineer to join our team. In this role, you will contribute to developing scalable internal data tools and modern data platform components. You will help build a robust data infrastructure that powers analytics across the organization.
Key Responsibilities
-
Collaborate in designing and implementing scalable data platform components, including Data Warehouse, Data Lake, and Feature Store
-
Participate in the deployment, automation, and monitoring of data platform infrastructure
-
Design, implement, and maintain efficient ETL data pipelines
-
Develop and support internal web-based tools for data access and operations
-
Work closely with data analysts, scientists, and engineers to understand data requirements and ensure high data quality
-
Continuously research and evaluate new tools, frameworks, and technologies related to data engineering
Requirements
-
Background in Computer Science, Data Engineering, or a related technical field
-
At least 2 years of experience in a data engineering or related role.
-
Hands-on experience with Apache Spark for distributed data processing
-
Familiarity with Apache Airflow for data workflow orchestration
-
Understanding of Hadoop ecosystem and large-scale data processing frameworks
-
Strong knowledge of Data Warehouse (e.g., star schema, snowflake schema) and Data Lake concepts
-
Basic understanding of RDBMS and NoSQL databases, and ability to write simple to moderately complex SQL queries
-
Familiarity with Unix environments, distributed computing, and version control systems (e.g., Git)
-
Familiarity with CI/CD pipelines, Docker, and infrastructure-as-code tools such as Terraform or Ansible
-
Basic experience developing internal web-based tools, with a focus on frontend development using ReactJS
Soft Skills:
-
Strong collaboration and problem-solving skills, with a proactive approach to working with cross-functional teams.
-
Ability to manage time effectively and prioritize tasks in a fast-paced environment.
-
Willingness to learn and adapt to new tools, technologies, and business needs
Nice to have:
-
Experience with cloud platforms such as AWS, GCP, or Azure
-
Knowledge of feature engineering and ML pipelines
-
Exposure to tools like dbt, Trino, Starrock, Datahub
-
Awareness of data governance and security best practices