Knowledge

PostgreSQL, Snowflake or Databricks: What a Mid-Sized Company Actually Needs

The question comes up in almost every first meeting, usually within the first half hour: “Do we need Databricks?” Sometimes it’s: “Our IT partner recommended Snowflake — is that a good choice?”

Our answer tends to be inconvenient, because it’s a counter-question: For what?

The tool question is almost always asked too early. Vendors sell platforms, and platforms sound like future-proofing. But a tool that’s bigger than your problem doesn’t solve your problem faster. It just costs more — and not only in money. It costs your data project its credibility when, a year in, an expensive platform sits there running three dashboards.

The three classes that matter

Class 1: the classic database. PostgreSQL is the best-known example — free to license, proven for decades, familiar to every IT service provider. What many people underestimate: a properly configured PostgreSQL instance handles tables with many millions of rows without breaking a sweat. For the order, customer and production data of a company with 50 to 250 employees, that is almost always more than enough. Running costs: hosting, a few hundred euros a month, often less.

Class 2: the cloud warehouse. Snowflake and Google BigQuery are the leading names here. They take over operations entirely, scale automatically and bill by usage. They make sense when data volumes grow, when many source systems converge, or when several sites need to run analyses in parallel. Pricing stays predictable as long as someone keeps an eye on query behaviour — unattended dashboards can inflate the bill.

Class 3: the data platform. Databricks is the heavyweight in this category: distributed computing across many machines, integrated machine-learning tooling, and the lakehouse concept that combines structured and unstructured data. Impressive technology — built for organisations with a dedicated data team to feed it, operate it and get value out of it.

Four questions decide

We use a simple decision path, and you can walk through it yourself.

Does your data fit on a single machine? As long as the answer is yes — and it is for the vast majority of mid-sized companies, including the ones that consider themselves data-rich — there is no technical case for distributed systems.

How many source systems feed in? One to three systems: Class 1 will do. From four to six systems with growing volumes, Class 2 becomes attractive, because operations otherwise turn into a permanent construction site.

Is there a data team? Without at least two or three people working on data full-time, Class 3 is off the table. Not because the technology is bad — but because nobody will be there to run it.

Are there ML ambitions beyond forecasting? Demand forecasts and early-warning systems run perfectly well on Classes 1 and 2. Only if you plan to train your own models at scale does a platform even enter the calculation.

The mistake we see most often

Buying too big. And, as an accelerant right behind it: the service provider’s home-grown framework. If your data preparation lives inside a custom system that only its creator understands, you have signed up with a vendor for life. That’s why we insist on standard tooling — above all dbt for data transformation, which by now virtually every data freelancer knows.

How error-prone the status quo is without such tooling has been quantified repeatedly: researcher Raymond Panko found over the years that roughly 88 percent of complex spreadsheets contain material errors; a 2024 literature review led by Pak-Lok Poon put the figure at 94 percent of business spreadsheets examined. The most famous single case remains the Reinhart–Rogoff spreadsheet, where a few omitted rows influenced global debt policy for years. If your reporting runs on chains of spreadsheets, these numbers are worth knowing.

The good news: the decision isn’t final

This is the part that takes the fear out of the tool question. If your data logic lives in dbt and SQL, it is portable. Moving from PostgreSQL to Snowflake then becomes a manageable infrastructure project, not a restart. Which means you are allowed to start small — and that’s exactly what we recommend: Class 1 or 2 depending on size, standard tools, one concrete use case. You can always grow.

For a quick self-assessment: our Data Readiness Check places you in five minutes — including an honest answer to the Databricks question.

Let's find out where AI can make the biggest difference for your business.

A no-obligation introductory call – we listen, ask questions, and honestly assess how we can help.

Schedule a free consultation →
30 minutes
Free of charge
Confidential