Big data projects rarely fail on tooling. They fail because nobody drew the layers first, so the team ends up with nine products that overlap in three places and leave a hole in the fourth.
A big data architecture is the drawing that prevents that. It says where data lands, where it rests, how it gets processed, who serves it to the business, and what keeps the whole thing running on schedule. Six layers, each with one job.
What follows is the practical version. You get what belongs in each layer, which tools actually sit there, how lambda, kappa and lakehouse differ and when each is the wrong call, what these decisions do to your monthly bill, and what the whole thing looks like once it is running in production.
Key takeaways
- A big data architecture has six layers: ingestion, storage, batch processing, stream processing, serving, and orchestration. Every component in the stack does one of those six jobs.
- Choose the pattern before the tools. Lambda runs batch and streaming side by side, kappa runs a single streaming path, lakehouse puts an open table format over cheap object storage.
- Jay Kreps made the case against lambda in 2014 and it still holds: keeping two code paths that must produce identical results is the most expensive part of the design.
- Storage is cheap and compute is not. Your bill is set by how often you reprocess data, not by how much of it you keep.
- If your team cannot name the layer a tool belongs to, you do not have an architecture. You have a shopping list.
The six layers of a big data architecture
Vendors group big data architecture components differently, but every stack, whether it runs on three services or thirty, resolves into the same six layers. The tools change. The jobs do not. Read the table first, then the detail underneath, and if you are still shortlisting big data architecture tools, work out which tool fits each layer before you sit through a single demo.
| Layer | Job | Common tools | What breaks without it |
|---|---|---|---|
| 1. Sources and ingestion | Get data in, from databases, apps, logs, devices and files | Kafka, Kinesis, Fivetran, Debezium | Teams write one-off scripts per source and nobody knows what is loaded |
| 2. Storage | Hold raw and refined data cheaply and durably | S3, ADLS, GCS with Iceberg or Delta | Data is trapped inside whichever engine wrote it |
| 3. Batch processing | Transform large volumes on a schedule | Spark, dbt, Hive | Reporting logic hides inside dashboards and drifts |
| 4. Stream processing | Act on events within seconds | Flink, Spark Structured Streaming, Kafka Streams | Every answer is at least a day old |
| 5. Serving and analytics | Answer queries fast enough for people and applications | Snowflake, BigQuery, Redshift, Druid | Analysts query raw files and time out |
| 6. Orchestration and governance | Run jobs in order, track lineage, control access | Airflow, Dagster, Unity Catalog | Silent failures, and nobody can say where a number came from |
Layer 1: Sources and ingestion
This layer moves data from wherever it is created into somewhere you control. It handles two shapes of traffic that behave nothing alike. Batch loads pull a table or a file drop on a schedule. Streams carry a continuous flow of events with no natural end.
Kafka is the default for the streaming half. It stores events in partitioned topics and keeps them for a retention window you set per topic rather than deleting them once a consumer reads them, which is what makes replay possible later (see the Apache Kafka introduction). For the batch half, change data capture tools such as Debezium read the database log instead of querying the table, so you get every change without hammering production.
The decision that shapes everything downstream happens here. What shape your source data is in, rows out of a database versus PDFs and support tickets, determines the storage format you can use and the processing engines that can read it. Most of the pain at this layer comes from the number of sources rather than their size, which is the ground the data integration challenges, solutions and tools guide covers in detail.
Layer 2: Storage
A big data storage architecture is now two things: object storage plus a table format. The object store, S3 or ADLS or GCS, holds the files. The table format sitting on top turns those files into something that behaves like a database table.
That second part is what changed in the last few years. Apache Iceberg and Delta Lake track snapshots, handle schema changes and let several engines read the same files without copying them. You can query one Iceberg table from Spark, Trino and Flink at the same time. Before table formats, each engine wanted its own copy, and you paid for storage three times.
Now we’re increasingly seeing a model where storage is an object store. It’s not a local file system; it’s a remote service, and it’s already replicated. Building on top of an abstraction like object storage lets you do fundamentally different things compared to local disk storage.

This is also where the warehouse versus lake question gets settled, and for most teams the answer now is both. The trade-offs between data warehousing and data lake architecture are worth working through before you commit, because the format you choose here locks in every engine above it.
Layer 3: Batch processing
Batch handles the work that is too heavy to do on arrival. Joining a year of orders against a customer table, rebuilding a model input, recalculating attribution across every channel. It runs on a schedule, usually hourly or nightly.
Spark still does the heavy lifting for large joins and file-level work. dbt has taken over the SQL transformation layer because it puts version control and tests around logic that used to live in stored procedures nobody could find. Most teams run both, with Spark handling the raw-to-refined step and dbt handling refined-to-reporting.
Layer 4: Stream processing
Stream processing is where a real time big data analytics architecture actually lives. It runs continuously and produces answers in seconds. Fraud scoring, inventory counts, alerting on sensor readings, anything where a nightly batch would deliver the answer after it stopped being useful.
Flink is the strongest option when you need real event-time handling, watermarks and exactly-once guarantees across a long-running job. Spark Structured Streaming is a reasonable choice if your team already runs Spark and your latency target is seconds rather than milliseconds. Kafka Streams fits when the processing is simple and you would rather not run a second cluster.
I would argue that well-designed streaming systems actually provide a strict superset of batch functionality.

Capability is not the same as need, though. The honest test for this layer is whether anyone acts on the output within minutes. If the answer is no, you are paying streaming prices for batch value.
Layer 5: Serving and analytics
The serving layer holds data in a shape that answers questions quickly. Analysts querying raw Parquet across a lake will wait minutes. The same query against a columnar warehouse returns in seconds because the data is sorted, clustered and often pre-aggregated.
Snowflake, BigQuery and Redshift cover general analytics. Druid, Pinot and ClickHouse cover the narrower case where an application needs sub-second responses on high-cardinality data. Once this layer outgrows a single database it stops being a storage decision and becomes a modeling one.
Read more: The 6 Components of an Enterprise Data Warehouse (EDW)
Layer 6: Orchestration and governance
Nothing above works without something deciding what runs, in what order, and what happens when a step fails. Airflow remains the common answer. It defines pipelines as DAGs, so dependencies and scheduling live in code rather than in a cron file (Airflow DAG documentation). Dagster takes a data-asset view instead, which suits teams who think in tables rather than tasks. A scheduler tells you a job ran, not that its output was right, and data observability for ETL pipelines is how teams close that gap.
Governance belongs in the same layer because it answers the same class of question. Who can read this, where did this column come from, and which downstream report breaks if we change it. Bolting it on afterwards is the most common reason a pipeline gets rebuilt, so it belongs inside the data engineering services scope from the first sprint rather than a phase two.
Big data platform architecture, or a stack you assemble yourself
There are two ways to fill the six layers. A big data platform architecture buys most of them from one vendor, so Databricks, Microsoft Fabric or Snowflake covers storage, processing and serving under a single contract. The alternative is a big data infrastructure architecture you assemble yourself, picking the strongest option at each layer and wiring them together.
Platforms win on time to first result, and on having one support number when a job dies at 2am. Assembled stacks win on cost at scale, and on not being repriced by a vendor who knows how much of your business runs on them. Team size usually settles it. In our experience, teams with fewer than five data engineers come out ahead on a platform once salaries are counted against license fees.
Lambda, kappa and lakehouse, and how to choose between them
The six layers describe what exists. A big data architecture framework describes how layers 3 and 4 relate to each other, and there are three worth knowing. The choice between them drives most of your cost and nearly all of your maintenance pain.
Lambda runs batch and streaming as two separate paths over the same data. The batch path is slow and correct. The speed path is fast and approximate. A serving layer merges them. It works, and it was the standard answer for years, but you write the same business logic twice in two different systems.
That cost is the reason kappa exists. Jay Kreps argued in 2014 that the reprocessing problem lambda solves is real, while the fix is worse than the disease. His alternative was one streaming path, replayed at higher parallelism when the logic changes. The full argument is in Questioning the Lambda Architecture, and it is worth reading before anyone on your team proposes lambda.
The problem with the Lambda Architecture is that maintaining code that needs to produce the same result in two complex distributed systems is exactly as painful as it seems like it would be.

Lakehouse sidesteps the argument. An open table format over object storage gives batch and streaming writers the same tables with transactional guarantees, so there are not two copies to reconcile. That removes the problem lambda and kappa both spend effort managing, and it is why a modern data lake services build and a warehouse build now look more alike than they did five years ago.
How to choose, in the order the questions actually matter:
- Does anyone act on data within minutes? If no, skip streaming entirely and run batch on a lakehouse. This is the most common right answer and the one teams talk themselves out of.
- Do you need both low latency and full historical accuracy on the same metric? That is the narrow case where lambda is still worth its cost, usually in finance and ad tech.
- Is your team small? Kappa on a managed streaming service is cheaper to run than two paths, but only if you have someone who understands event-time semantics.
- Are multiple engines reading the same data? Then the table format matters more than the pattern, and lakehouse is the call.
Sequence the build the same way. A big data architecture strategy should deliver in this order: prove the ingestion and storage layers with one real source, get a query in front of an analyst, then add processing complexity. Teams that design all six layers on a whiteboard before landing a single row usually spend two quarters before anyone sees a number. Data engineering in Microsoft Fabric walks through the same six layers assembled inside one vendor platform, which is a useful comparison if you are weighing that against a best-of-breed stack.
What a big data architecture costs to run
Vendor pages describe a big data solution architecture and never mention what it costs to run. Here is the shape of the bill.
Storage is the cheap part, and it is rarely what breaks a budget. Google Cloud Storage pricing lists Standard storage in a single US region at $0.020 per GB per month, and the other major clouds sit in the same range, so ten terabytes of raw data costs roughly $200 a month to keep.
Compute is where the money goes, and the driver is not volume, it is frequency. A transformation that runs hourly uses roughly twenty-four times the compute of the same transformation run nightly. Streaming costs more again because the cluster never stops. A modest Flink or Kafka Streams job running continuously will outspend a nightly Spark job on the same data by a wide margin, so the latency target you set in layer 4 is the largest line item you actually control.
The third cost is the one nobody forecasts: people. A lambda architecture needs engineers who can maintain two implementations of every rule. That is a permanent headcount line, and it is usually larger than the infrastructure it supports.
What a big data application architecture looks like in practice
Contract analysis is a useful workload to trace through the layers, because the input is messy and the output has to be right. Here is how the layers map, as a pattern rather than a spec.
Contracts arrive as PDFs, scans and even handwritten pages, which means layer 1 is a file-drop watcher rather than a database connector. Layer 2 keeps the original documents untouched and writes the extracted terms to a table, so the source of truth is never overwritten. Layer 3 runs the extraction in batch, because a contract signed on Tuesday does not need reading by Tuesday afternoon. There is no layer 4 at all, and that absence is the design decision, not an oversight. Layer 5 serves the extracted terms to the systems that bill against them. Layer 6 records which model produced which result, which is the first thing an auditor asks for.
The AI contract analysis case study shows the same shape at scale: more than 50,000 contracts averaging 15 to 20 pages, three models running in parallel and cross-checking each other, and extracted billing terms flowing straight into account configuration at 95% extraction accuracy.
Two things are worth copying from that shape. The architecture skips a whole layer on purpose, and the hardest requirement, accuracy that finance can trust, came from the business rather than from the data team.
How Brickclay helps
We build these six layers as one system rather than six procurements. The first delivery is always a working slice: one real source, landed in storage, transformed once, and queryable by an analyst. Nothing gets standardized until that slice runs on production data, because a layer that looks fine on a diagram tends to show its problems the first time real volume moves through it.
If you are deciding between lambda, kappa and a lakehouse, or you have a stack that grew tool by tool and now nobody can say which layer owns what, that is the point to bring in big data services. Map what exists against what the six layers require, then decide which parts are worth keeping.
Work with Brickclay
Whatever you just read about, we build it.
Brickclay is a digital transformation partner with multiple disciplines in one team: data and analytics, AI and automation, cloud infrastructure, product engineering, brand experience and digital marketing. 100+ specialists. 300+ projects.
Tell us what you're building. We'll tell you which of our teams you need, and which you don't.