Data and analytics
Senior Data Engineer - Azure, Databricks & ML Pipelines | Remote - remote job
Toptal
Senior Data Engineer - Azure, Databricks & ML Pipelines | Remote - Data and analytics, remote
Toptal is hiring a Senior Data Engineer - Azure, Databricks & ML Pipelines | Remote - a senior role, so the employer expects you to make the calls yourself. The role sits in Data and analytics and is fully remote. The company places no restriction on where the candidate lives, so you can apply from anywhere.
The main tools named in the advert are: Python, SQL, Machine Learning. Your CV should show concrete examples of exactly these skills.
The employer did not publish a figure; that is settled at interview. The advert sets no condition on working hours.
This vacancy passed an automated check: the list excludes any advert requiring a foreign work permit, visa sponsorship, a particular citizenship, or residence in a specific country.
In brief
- Company
- Toptal
- Category
- Data and analytics
- Who may apply
- From anywhere in the world
- Work mode
- Fully remote
- Level
- Senior
- Tools
- Python, SQL, Machine Learning
- Posted
- 29 July 2026 (27 days ago)
- Active
- until 7 September 2026
- Source
- We Work Remotely
Get new Data and analytics jobs by email
The board updates several times a day. We write only when a new, checked vacancy appears. No spam, no account needed.
Already subscribed? Manage your settings
The employer's description
Headquarters: Remote
URL: https://www.toptal.com/
About the Role
We're looking for a Senior Data Engineer to design, build, and maintain scalable data pipelines and ML-ready infrastructure on Azure and Databricks. This is a hands-on engineering role: you'll own the full data pipeline lifecycle - ingestion, transformation, orchestration, and deployment - while supporting machine learning workflows with clean, reliable data. If you're comfortable owning infrastructure decisions and writing production-quality Python at scale, this role is built for that.
What You'll Do
* Design, build, and maintain data pipelines using Databricks and Azure-native data services
* Develop and optimize ETL/ELT processes to support analytics and machine learning workloads
* Build and maintain CI/CD pipelines for data engineering and ML deployment workflows
* Write clean, efficient, production-quality Python for data processing and pipeline automation
* Support machine learning teams with well-structured, high-quality datasets and feature pipelines
* Design and manage data architecture across Azure services (e.g., Azure Data Factory, Azure Data Lake, Azure Synapse)
* Monitor pipeline performance, troubleshoot data quality issues, and implement reliability improvements
* Implement data governance, security, and access control best practices
* Collaborate with data scientists, analysts, and software engineers to align data infrastructure with business needs
* Participate in code reviews, architecture discussions, and technical planning
What You Bring
* Strong hands-on experience with Azure cloud data services
* Proven experience building and maintaining pipelines on Databricks
* Solid experience designing and managing CI/CD pipelines for data or ML workflows
* Strong Python skills for data engineering and pipeline development
* Working knowledge of machine learning workflows and how data engineering supports them
* Experience with SQL and relational/distributed data systems
* Understanding of data pipeline orchestration, monitoring, and reliability practices
* Strong problem-solving skills and ability to work independently on complex data infrastructure challenges
* Solid communication skills for collaborating with data science and engineering teams
Nice to Have
* Experience with MLOps practices and tools (MLflow, Azure ML)
* Familiarity with Spark internals and performance tuning within Databricks
* Experience with infrastructure-as-code (Terraform, Bicep, ARM templates)
* Exposure to real-time/streaming data pipelines (Kafka, Event Hubs, Structured Streaming)
* Relevant Azure or Databricks certifications
Why This Role
* Full pipeline ownership: Own data infrastructure end to end, from ingestion through ML-ready delivery
* Modern data stack: Work with Azure and Databricks, leading platforms in enterprise data engineering
* Cross-functional impact: Directly enable machine learning and analytics outcomes, not just move data
* Flexibility: Remote-friendly engagement structure
How to Apply
Ready to bring your data engineering expertise to Azure and Databricks-powered ML infrastructure? Apply through Toptal here: https://www.toptal.com/talent/apply
To apply: https://weworkremotely.com/remote-jobs/toptal-senior-data-engineer-azure-databricks-ml-pipelines-remote
The text is kept in the employer's original language, because that is the language you will apply in.
Frequently asked questions about this job
Can I apply for Senior Data Engineer - Azure, Databricks & ML Pipelines | Remote from where I live?
Yes. Toptal accepts candidates for this vacancy from anywhere in the world, so you need no work permit for a foreign country. The advert passed an automated check: had the employer required a work permit, visa sponsorship or residence in a specific country, it would not be on this board.
What pay is stated?
Toptal did not publish pay for this vacancy. Most remote adverts publish no figure; it is settled at interview.
How do I apply?
You apply directly to the employer, through the original advert published on We Work Remotely. Donator takes no applications, charges no commission and stores no CV.
What kind of vacancy is this?
A fully remote role in Data and analytics. Hybrid adverts and anything requiring office attendance are not published on this board.
Free certificates for this vacancy
This advert asks for Python, SQL, Machine Learning. Below are the free credentials that cover exactly those tools.
Kaggle Learn: 17 micro-courses
certificateKaggle · ~4 h
Intro to ML, Pandas, SQL, Deep Learning, Computer Vision and more. Only a Google account is needed, and the certificate has a public link.
moderate weightCS50x: Introduction to Computer Science
certificateHarvard CS50 · ~100 h
Harvard's own branded certificate is free once you pass every problem set and the final project. The edX verified certificate is a separate paid product and is not needed.
strong brandDeveloper certificates (10+ tracks)
certificatefreeCodeCamp · ~300 h
A non-profit; a card is never requested. The certificate is issued after five projects are submitted and pass.
moderate weightDeep Reinforcement Learning course
certificateHugging Face · ~30 h
Entirely free with no deadline. 80% earns a Certificate of Completion, 100% a Certificate of Honors.
moderate weight
Similar jobs
Senior ML Engineer (Europe-based/Remote)
SWORD Health · 4 days ago
From EuropeMachine LearningData Analytics Engineer
Ruby Labs · 4 days ago
From EuropeSQLData Scientist (AI Data & LLM Specialist)
Eclipse Laboratories · 5 days ago
From anywherePythonMachine LearningSenior Data Analyst
Ruby Labs · 5 days ago
From Europe