Are Data Teams Still using Airflow For Their Data Pipelines?

Are Data Teams Still using Airflow For Their Data Pipelines?

October 6, 2026 data engineering 0
airflow tutorial

Someone once called me the “Airflow Guy“. I found it a little strange. I have really only written a few articles on Airflow. But you never know what sticks with people.

Airflow was a very popular data orchestration solution in 2020. It was only about five years old at the time.

Fast forward a few years and suddenly there are plenty more open-source orchestration options. So will data teams still be using Airflow in 2027?

Is Airflow the solution you should build your data workflows on?

In this article, I wanted to share what I am seeing in terms of data solutions and platforms and where they are going.

Data Workflows

If you’ve worked in data even for a few years, you’ve likely at least heard about Airflow. It’s probably one of those tools on a list that someone recommends you learn. There is a reason for that.

When Airflow was released, many companies and data teams were still building their own custom data orchestration setups. For many, Airflow was…a breath of fresh air. Jokes aside.

Of course there were other options. You may have seen or heard of Apache NiFi or SSIS. Not orchestrators, but many data teams use these solutions somewhat interchangeably. You did have to often pair them with other components such as SQL Server Agent for SSIS to schedule jobs, but it worked.

Airflow was a Python framework. Really, the perfect language because everyone was already starting to adopt Python heavily for data work. It was easy to adopt. It also made making data pipelines easy.

So many data teams switched.

Everyone Wanted To Replace Airflow

Airflow quickly grew, but several other solutions grew alongside it. Dagster, Prefect, Mage, Orchestra, to name just a few.

For a while, nearly every new Python orchestration framework(and some GUI based solutions) seemed to get compared against Airflow.

Each of them had some edge they leaned onto.

Better developer experience. Better local testing. Better data asset tracking. Better observability. In a funny way, Airflow became the benchmark everyone sold against.

All positioning themselves as Airflow replacements back in 2020. Now many are switching, but back then, Airflow was the target. More of the market was likely dominated by other data pipeline solutions; however, when it came to open-source Python solutions.

There were only a few.

They all wanted to replace Airflow.

It’s funny because most of them barely talk about Airflow anymore. Some are more focused on AI/ML workflows; others have merged or been acquired.

And I think that shift says something interesting about the orchestration market.

Airflow wasn’t necessarily “beaten.“

The problem that data teams were trying to solve started changing.

The Problem Changed

Many original data orchestration solutions were simply stored procedures wrapped in a bash or Python script called by Cron or Jenkins. 

You had APIs you needed to call, transforms that would run over data, and a few reports to support.

Back then, Airflow made a lot of sense. Especially since many data platforms really only focused on being a data warehouse. Now more of these solutions are providing ways to manage and schedule jobs internally, even with third-party solutions. I am half surprised Snowflake hasn’t provided an Airflow-managed solution or at least improved their own Tasks product so that it could compete more directly. 

Orchestration Isn’t The Same As Scheduling A Transformation

I did want to call out the fact that I do tend to use data pipelines and orchestration interchangeably. These are different, but functionally, on the data side, many data teams rely on Airflow and other orchestrators similar to the way they do data pipeline solutiosn.

Scheduling a dbt model every hour is not necessarily the same problem as orchestrating a workflow across five different systems and not just data systems.

Imagine one data team whose workflow looks like this:

If all of that is happening inside Snowflake or Databricks, you probably don’t need a separate platform coordinating every step. Snowflake Tasks, Databricks Workflows, dbt, or whatever platform-native scheduler you’re already using might be enough.

Now imagine another workflow:

Simply putting dbt over all of this will prove far more difficult. Don’t get me wrong, you could set-up some Python transforms that maybe fill in some of these gaps. But it’d likely be a pretty complex workaround.

This is also why some data teams can use an EL solution and dbt on top of Snowflake and others might have to use a more holistic solution.

Not Every Data Team Needs Airflow

When you actually look at companies’ data infrastructure today.

I’ll start with the fact that I still run across plenty of Airflow; it has its place.

But one thing I’ve noticed is that many companies are trying to limit how many solutions to use. If they are building on Snowflake, they likely are just using an ELT tool like dbt paired with Snowflake-hosted dbt. It’s light, easy to set up, and will do most everything your data team needs in terms of transforms.

This is actually my go-to in many cases. If you need to pull from databases like MySQL and Postgres as well as ingest data from tools like NetSuite and Salesforce, Estuary has been a reliable choice. I do always want to be upfront and say I advise them, so you can take that line with a grain of salt.

For data teams using Databricks, they might have a similar set-up. Databricks would prefer you use all of their own tooling like Databricks Lakeflow, but I do provide a caveat here to clients. If you build using any internal solutions, it’ll be that much harder to migrate away. So I still recommend people use something like dbt. It’s easy to move around and gives you more leverage when negotiating contract renewals.

Can you imagine how much leverage a company can have over you if the entire data stack is one solution? I’ve seen some data platforms be ruthless in contract negotiations, especially at the enterprise level.

Airflow’s Biggest Advantage Might Be That It Is Boring

There is also something Airflow has that newer tools can’t easily recreate: maturity. We know the sharp edges and what we dislike about Airflow.

I’ve spoken to a few teams who have used the new solutions and felt like Airflow might have met their needs better. Now the grass is always greener on the other side. But most of them picked those alternatives to Airflow because of all the fear-mongering around Airflow.

In many ways, those lists of problems that Airflow may have at least let you know what the cons are. With new or lightly used solutions, you really are limited in knowing their gotchas. 

Airflow has a strong user base, lots of documentation, known workarounds, etc. So in many cases, that makes Airflow a reliable choice when it comes to picking your solution.

New alternative orchestration and data pipeline tools always know exactly how to sell against Airflow’s weaknesses. They specifically target marketing to make it sound like they are superior. At the same time, they ignore their own trade-offs.

This doesn’t mean you shouldn’t seriously consider Airflow’s trade-offs. 

Airflow’s Biggest Weakness Is Also Airflow

On the other hand, Airflow can become a lot of infrastructure to operate.

Anyone who has worked in a large Airflow environment probably knows what I mean. When you first start out using Airflow. It feels pretty lightweight. 

Just run:

airflow standalone

Suddenly you have an orchestrator running locally with a UI and a scheduler. 

You go and build a dozen or so DAGs, and everything works great, until it doesn’t. It usually starts with some DAGs getting backed up because you’re likely using the Sequential Scheduler, even though they tell you not to. It works, right?

Then you run into space issues with logs, or because you never switched from using the SQLite instance Airflow sets up.

And that’s all assuming you’re building by yourself. You don’t have an entire data team building with you. Add a group of 10 more data engineers, and things are going to get messy fast.

Airflow Probably Isn’t Going Anywhere

I don’t think Airflow is going anywhere. Data solutions stick around for a long time. And alternative “killer” versions of those tools rarely kill said solutions.

I remember when Domo was pitched as a Tableau killer. Now it’s barely worth $158M.

So do be careful when salespeople and articles tell you that a solution will replace what others might call industry standard. The truth is, if a company only has a few million in funding and is trying to replace a data tool that’s been around for over a decade and has had 100x the amount of time, effort, and funding, it’s got a long road ahead of it.

Most of the competitors try to do too much and don’t focus on the edge or features that really set them apart, so they are never good enough to truly compete. 

But perhaps I am starting to write a whole new article.

If you’d like to read more about data engineering and data science, check out the articles below!

The Data Engineer’s Guide to ETL Alternatives

Does ELT vs. ETL Even Still Matter?

The Data Engineering Job No One Wants To Do – Backfilling

Why Data Pipelines Exist – Beyond Moving Data From Point A To B

What Leading a Data Team Actually Looks Like Right Now

Schema Drift in Snowflake Pipelines and How to Handle It

How To Set Up Your Data Stack For 2026 – Data Infrastructure For AI

 

Leave a Reply

Your email address will not be published. Required fields are marked *