In 2026 The Data Fundamentals Matter More Than Ever
We are headed towards a massive data mess otherwise
Hi, fellow future and current Data Leaders; Ben here 👋
It’s 2026, and everyone wants to talk about AI engineers, FDEs, and agents.
But in almost every company I work with, the real bottleneck is still the same thing it's often been…messy data, weak data foundations, and teams that need help making data valuable.
Now before we jump in to talking about the data world in 2026, I wanted to share a bit about Estuary, a platform I’ve used to help make clients’ data workflows easier and am an adviser for. Estuary helps teams easily move data in real-time or on a schedule, from databases and SaaS apps to data lakes and warehouses, empowering data leaders to focus on strategy and impact rather than getting bogged down by infrastructure challenges. If you want to simplify your data workflows, check them out today.
Now let’s jump into the article!
When I first started in the data world, the sexiest job of the 21st century was a data scientist.
Of course, that’s what I wanted to do!
Even when it came to consulting, that’s what I was looking for. Data science projects!
But you know what I found out...every data science project was +80 percent data engineering. And really, the skills that got you far were really solid fundamentals.
In 2015, if you knew SQL, Python, and data modeling well, you could get a job at a lot of companies. Even Facebook’s interview was really just SQL, Python, and data modeling for the most part.
But it’s starting to feel like what is being demanded of new engineers and analysts is even more than that as if the world is shifting away.
The coolest job?
It’s somewhere between AI Engineer and FDE. And that requires a whole new set of skills.
But does it?
I wanted to share some of my thoughts in terms of what I am seeing in the data world in 2026 and what you should do to break in.
Even if the job titles change, many of the fundamental skills and needs of the business don’t.
The Basics Will Get You Far - It’s All About Your Foundations!!!
If we roll back time by ten to fifteen years, and you were wanting to learn the basics as a programmer and developer, you likely picked up a lot of ancillary skills you didn’t even realize.
Spinning up Docker containers, ingesting CSV files with wonky formats and jagged rows, working with SFTP or other ancient technologies. All so you could set up Airflow.
Ok, maybe not Airflow, but you’d likely find yourself having to do a lot of pre-work.
In turn, you’d pick up what I like to call “glue technical skills”. The stuff that’s in between running data pipelines and data warehouses. The stuff that actually makes everything run.
The stuff you assume everyone knows, but no one who hasn’t had to struggle with it for four hours on a weekend knows. It’s also likely what will give you a good sense when AI gives you garbage vs something that looks like it’s going in the right direction.
The basic skills you pick up are your foundation(I should also add that data modeling is also critical, perhaps that is another future article).
Sure, they might not always be the coolest skills. But they are often the most useful. They let you understand how to build more reliable AI systems and workflows. In fact, many of the skills I see people incorporate into these systems…are traditional programming and system design in one way or another.
I do tend to have a bias here. I went through a similar hype cycle in the culinary industry. And as one of my chefs used to say:
“You know what that dish is missing for garnish, a little solid technique”
Maybe it’s also why I recently wrote the post below:
Because sure, we might, in the future, get further abstracted away from some of the work we are doing today. But I don’t think it will ever completely disappear. The same way, to this day, there is still a need to write in lower-level languages.
Here We Go Again
Along with the basics still being relevant, here is another thing I still see as relevant.
Or at least, the act of centralizing data in a single location and making it more accessible to multiple users(and now machines) to help improve performance, reduce costs, and create a more consistent set of naming and entities.
Don’t get me wrong. I am sure there are plenty of people saying, “ Hey, just leave your data in the source system or in some very raw state, and then have your AI go over that.
Yeah, we did that in 2010; it was called schema on read. It went really poorly. It’ll also likely drive your token costs through the roof. So the labs will be happy.
I am not saying a data warehouse or lake house is the answer to everything. There are plenty of people who are doing just fine on reporting on database replicas. But as your company grows and has increasing data and complexity. It is often a good choice.
And even now, in the era of AI. I am seeing the same pattern.
Companies’ data is disparate, and they feel limited on the questions they can answer, so they don’t dig too deeply since everyone is busy.
Then, someone finally does centralize their data, and then the floodgates are open. The challenge then becomes focusing on the right use cases, ensuring there is some governance. Because if you don’t, then you start having the same dashboard sprawl, or perhaps in this LLM-fueled present, agent sprawl that will cause minor data inconsistencies that will blow up to indecision and push back from leadership.
How Can You Break Into Data In 2026?
I imagine if you’re just starting your journey in the data world, it feels overwhelming.
I was talking to someone looking to grow as a data engineer, and they brought up the fact that they were scrolling through the data engineering subreddit and that they didn’t even understand all the technologies being mentioned.
Mind you, this was someone who had built out their own data warehouse, an analytics agent to ask questions of said warehouse, data pipelines, along with a few one-off Python connectors.
Back in the day, most of us were probably lucky to write SQL scripts and SSIS packages on someone else’s data warehouse because fewer companies were building them. You couldn’t just spin up a cloud data warehouse back then.
Someone had to get space on a server or purchase a server(or there is the classic Laptop server). Then you’d be able to spin it up….wait, someone needs to set it up on the network, set the correct firewall rules, and so on.
Now, go sign up for some trial somewhere, and you can start ingesting data very fast.
Then, ask an LLM to develop some raw data for you, or find a free API, and start building your pipelines.
Yes, it’s not building in production.
But the barrier to starting to get better at data engineering has never been lower. The tools are all there. Many are free.
In fact, if you are on LinkedIn, almost every data engineering creator will post nearly the same post about the topic every 3-6 months(I know I have at some point!).
So here is what I would do.
Get really good at the basics - Meaning you should be able to have a conversation about data modeling, SQL, and software design principles, and maybe get a few opinions on how you think its impacted by AI. What do you think should change? Why?
Get comfortable with messy data - I believe data will get even messier in the future. Engineers are being forced to move faster. Meaning there are a lot of systems being developed more poorly than in the past. There will likely be missing ID fields, update and create dates, poor integrations across systems, and that’s just a few initial thoughts. Be ready for a world where you will spend a lot of time integrating and parsing semi-structured JSON data sets that were built only to function in a specific application and never be parsed and analyzed.
Think about where AI could really impact the work you do! It’s still so early in this LLM era. Everyone is learning exactly how LLMs can be useful. You might actually be farther ahead than you think. Learn the basics, then consider how LLMs could be used to improve key data tasks. Ask yourself, how could a migration workflow be improved? What can you do differently today that we couldn’t even consider a decade ago?
Build something end-to-end - Tutorials are a trap. Build something end-to-end. Instead of just rewatching how to set up Airflow for the tenth time, go out and build an end-to-end example. Create a basic front-end for it. Have fun, D3.js is free, Tableau has a trial, and you can even try building a full-blown website.
Final Thoughts
The business world always wants some magic solution to fix its data and business problems.
Just become data driven.
Use self-service analytics.
Be AI-native.
These all sound nice. They make it sound like, hey tomorrow, if you buy our tool you will be exactly where you want to be. You won’t. There is plenty of work to be done to get your business from where you are now, to “AI-Native” and just purchasing a subscription to Claude isn’t the answer.
The same goes for you reading this. You also have to work on your basic skills. Maybe you will write less of your own code in the future, but you’ll still need to implement what the AI builds. It’s why I like asking interviewees about old problems and how they’d approach it today.
How would you run a migration today?
Could you incorporate AI? Which problems do you think it’d solve best or not?
Where would AI likely perform poorly and how could you reduce any of the known issues?
Just because there is a new magic eight ball that can give you pretty okay answers, doesn’t mean you can’t think in the data world in 2026.
With that, thanks for reading!
Video Of The Week - If AI Can Replace Workers, Why Is It Hiring Consultants
Articles Worth Reading
There are thousands of new articles posted daily all over the web! I have spent a lot of time sifting through some of these articles as well as TechCrunch and companies tech blog and wanted to share some of my favorites!
Life Outside the Bay Area Bubble - Atoms, Bits, and the Resurgence of Detroit
By Joe Reis with Ryan Dolley
A question that comes up almost everywhere I go is, “Do you actually need to live in the Bay Area to do meaningful work in AI, tech, and data?” If you’re building the next foundation model or trying to raise a mountain of capital, San Francisco is undeniably the center of the universe. But for those building in the real world, the “atoms vs. bits” distinction is becoming a competitive advantage. While the coasts are hyper-fixated on pure software and AI tooling, places like Detroit are dealing in atoms (mobility, robotics, and heavy manufacturing), providing a “real-world grounding” that the Bay Area bubble often lacks.
The 5 Silent Failures in Data Pipelines
It’s 4:57 PM on a Friday, and a junior analyst noticed something odd.
The numbers in this week’s dashboard looked exactly like last week’s.
Not just similar.
Not kind of close.
Identical.
The pipeline had been rerunning but the data hadn’t changed for seven days, and nobody had a clue(until the CFO notices on a Monday morning).
Of course, maybe there is another story here about the fact that no one noticed that the data was stale, but let’s not get into that.
One of the challenges with data pipelines is that they can fail without anyone noticing.
Dashboards might only be looked at once a month.
Data can “look” right.
Pipelines can run without triggering any failure or red flag.
And what’s worse is that it can happen in multiple ways. In this article, I wanted to discuss the ways pipelines can fail silently and what you can do about it.
End Of Day 220
Thanks for checking out our community. We put out 4-5 Newsletters a month discussing data, tech, and start-ups.
If you enjoyed it, consider liking, sharing and helping this newsletter grow.





I have a feeling that it is not limited to data - the raise of AI actually breaks the barriers and moats made by fancy stuffs (for the mass, at least), and the real differentiator is just the basics: common sense, deep understanding of certain domains, judgements, etc.
I guess it is actually a good thing, that AI let us focus on what really matters and free us from low value stuffs.