Data Engineer

  • A termo incerto
  • Full time
  • €30,000 - €45,000
  • Remote
  • Data Engineering

About:
We're looking for a Data Engineer to own the design, implementation, and operation of production-ready data systems. You'll independently deliver end-to-end pipelines, from ingestion through transformation and serving, making the technical decisions that keep our systems reliable, scalable, and maintainable across the full data lifecycle.



What you will be doing:

- Design and implement end-to-end production data pipelines covering ingestion, transformation, orchestration, and serving layers.

- Build and operate distributed and non-distributed processing systems, choosing the right approach based on workload, data volume, latency, and cost.

- Design transformation layers with tools such as dbt, ensuring dependency management and CI/CD integration.

- Select the right storage technologies (relational and non-relational) for each system and workload.

- Design and maintain orchestrated workflows with frameworks such as Airflow or Prefect.

- Implement batching and incremental processing strategies to handle growing datasets efficiently.

- Build systems that behave reliably under retries, failures, and increasing scale.

- Ensure systems are observable and diagnosable through logging, monitoring, and validation.

- Work closely with DevOps to deploy and operate systems in production environments.

- Design and operate pipelines supporting AI and agentic workloads, considering freshness, semantic consistency, and point-in-time correctness.

- Take ownership of reliability and maintainability well beyond initial delivery.

- Participate actively in code reviews, identifying performance and reliability improvements.

- Support and mentor junior engineers in engineering practices and technical implementation.



What we're looking for:

- 3+ of experience in data engineering (a reference point: at JTA, competence matters more than years).

- Ability to independently deliver end-to-end pipelines across ingestion, transformation, orchestration, and serving.

- Strong Python skills for production-grade data processing: clean, tested, maintainable code.

- Strong T-SQL for building, exploring, validating, and troubleshooting datasets.

- Hands-on experience with dbt, including dependency management and CI/CD integration.

- Experience with code-first orchestration frameworks (Airflow, Prefect).

- Experience with distributed processing (e.g. PySpark), and sound judgment on when distributed vs. non-distributed processing is appropriate.

- Comfortable with both relational and non-relational databases, able to choose the right one per workload.

- Experience with containerized environments (Docker) and collaborating with DevOps on deployment and operations.

- Solid engineering practices: testing, logging, observability, code review, and version control (Git).

- Ability to investigate and resolve failures, inconsistencies, and reliability issues independently.

- Understanding of data governance, cataloging, and lineage principles.

- Ownership mindset: you handle ambiguity with autonomy and stay accountable for your systems beyond delivery.

- Team-oriented, proactive communicator who provides guidance and feedback to more junior engineers.

- Fluent in English, both written and spoken.

- Bachelor's degree in Computer Science, Software Engineering, Data Engineering, or a related field (or equivalent experience).

Nice to have: experience designing pipelines for both BI and AI workloads.

|
|
Powered by Factorial
Build my own jobs page