Medallion Architecture on Microsoft Fabric vs. AWS: A Practical Build Guide

Microsoft is calling Fabric its most ambitious data platform investment in a decade, and Databricks-style medallion architecture -- Bronze, Silver, Gold data layers -- has become the default reference pattern anyone researching a modern data platform runs into within the first five minutes. What gets lost in that conversation is that the platform you pick matters less than most vendor pitches imply, and the actual risk sits somewhere else entirely: Gartner's research, cited across LumenData, Oracle, and dozens of migration consultancies, puts data migration project failure -- projects that fail outright, blow their budget, or run past schedule -- at 83%. Choosing Microsoft Fabric over AWS, or the reverse, is a real decision with real cost implications. Building the thing badly is the decision that actually sinks projects.

Medallion architecture organizes data into three progressive layers. Bronze holds raw, unprocessed data landed exactly as the source system produced it -- append-only, full history retained, so nothing is ever lost to a bad transformation downstream. Silver is where data gets cleaned: deduplicated, type-validated, joined, and standardized, but still at a granular, atomic level rather than pre-aggregated for any one report. Gold is the business-ready layer -- dimensional models and aggregates built for a specific BI dashboard, ML feature store, or reporting need. Databricks popularized the terminology and the pattern around 2019-2020 while promoting its Delta Lake-based lakehouse, and it's now the de facto standard for lakehouses built on Delta Lake, with Delta Live Tables built explicitly around it. A 2025 serverless pipeline implementation that adopted the pattern reported up to 70% lower ETL latency while supporting near-real-time analytics on high-volume streams -- but the pattern has a well-documented failure mode of its own: creating too many Gold tables with no clear owner, which turns a clean three-layer model into an unmaintainable sprawl within a year or two.

Building this on Microsoft Fabric means starting with OneLake -- a single Delta-format storage layer (built on ADLS Gen2) that every Fabric engine reads from via shortcuts, so data exists once rather than getting copied between tools. Ingestion into Bronze runs through Data Factory pipelines or Dataflows Gen2; Bronze-to-Silver cleansing typically runs in Spark notebooks that write back to Delta tables; Silver-to-Gold modeling happens in notebooks or in a Fabric Warehouse using standard T-SQL. The payoff on the serving side is Power BI's Direct Lake mode, which reads Gold Delta tables straight out of OneLake with no import or duplication step. Pricing runs on capacity units: OneLake storage is billed at $0.023/GB per month, pay-as-you-go compute runs about $0.18 per CU-hour, and F-SKUs scale on a doubling ladder from F2 (roughly $260-310/month, a realistic small-business entry point) up to F64 (roughly $8,000-8,500/month on-demand, about 41% less with a one- or three-year reservation) and beyond. Microsoft has been named a Leader in Gartner's Magic Quadrant for Data Integration Tools for five consecutive years running, which is a reasonable proxy for how mature the tooling actually is.

Team presenting data charts and planning a project on a whiteboard in a modern office

The AWS-native equivalent swaps OneLake for S3 as the storage foundation, typically using Apache Iceberg or Delta Lake table format on top for the ACID transactions and schema evolution a lakehouse needs. Ingestion into Raw runs through Glue jobs and crawlers (crawlers bill at $0.44 per DPU-hour), AWS DMS for database change-data-capture, Kinesis Firehose for streaming, or AppFlow for SaaS sources. Raw-to-Silver cleansing runs in Glue ETL (PySpark) or EMR Spark jobs, writing Iceberg or Delta tables back to S3; Silver-to-Gold modeling either stays in Glue/EMR or loads into Redshift, either as native tables or queried externally through Redshift Spectrum. AWS Lake Formation handles governance -- permission-setting itself is free, though the underlying S3 and Glue Data Catalog usage bills separately -- and Step Functions or Managed Workflows for Apache Airflow (MWAA) orchestrate the whole pipeline end to end. Databricks-on-AWS is a legitimate hybrid path here too: Databricks scored at the top of the Lakehouse use case in Gartner's 2025 Magic Quadrant for Cloud Database Management Systems, and 2026 interoperability announcements mean Snowflake can now query data sitting in Microsoft OneLake directly through open standards -- genuine cross-platform architectures are viable in a way they weren't two years ago.

The question almost nobody asks before picking a platform is whether the business needs a medallion-pattern lakehouse at all. A data lakehouse earns its complexity when an organization is running both structured BI reporting and advanced analytics or AI workloads at the same time, and the friction between them is real and measurable -- financial analysts querying a warehouse, data scientists working against a separate data lake, duplicated pipelines, and weeks of engineering effort just to get one team access to what another team already has. The concrete signals are latency under real concurrency, queries routinely joining five or more historical tables, data sources multiplying across several source systems, or AI features that need genuinely fresh data pipelines rather than a nightly batch. Absent those specific pain points, a small or mid-sized business does not need this complexity yet, and the better move is building an analytics program in small, incremental steps rather than standing up a full Bronze-Silver-Gold platform because a vendor pitch or a conference talk made it sound inevitable.

Gartner's 83% failure rate isn't abstract, and it isn't really about picking the wrong cloud. The Bloor Group's research on the same class of projects found average cost overruns of 30% and average schedule overruns of 41%, and one of the most consistently underestimated line items is data transfer -- egress fees alone typically run 6-12% of total migration cost, and most project budgets don't account for them until the first bill arrives. For a small business migrating a genuinely limited workload, realistic project costs run $10,000-$50,000; for a business that didn't actually need a full lakehouse build and got sold one anyway, that number climbs fast, and the project still carries the same odds of blowing its budget, missing its schedule, or simply not shipping. The platform choice is a real decision. The decision to build a full enterprise data platform in the first place, before confirming the business has the access-friction and AI-pipeline problems that pattern is designed to solve, is the one that actually determines whether a project lands in that 83%.

Whichever platform gets chosen, the build sequence is the same, and it's worth being explicit about where the engineering time actually goes. Land source data unchanged into Bronze/Raw first, with a separate ingestion path per source system. Bronze-to-Silver is where most of the real work happens -- deduplication, schema validation, and MERGE-based upsert logic to handle records that arrive late or get updated after the fact, plus slowly-changing-dimension handling for anything that needs historical tracking. Silver-to-Gold is modeling: star schemas, aggregates, or feature tables built for whatever's actually going to consume them, whether that's a BI tool or an ML pipeline. Orchestrate the whole thing end to end with monitoring and alerting so a failed job gets caught before a dashboard silently goes stale. Governance and access control come last in the build sequence but should never be an afterthought in the plan -- if PHI, PII, or other regulated data flows through any layer of this pipeline, the same access-logging and encryption discipline that applies to a HIPAA-regulated network applies here too, and it's far cheaper to design in from the start than retrofit after an audit finds the gap.

Mytek Pros isn't a Spark or ETL development shop, and we won't pretend to be one -- but the decision that precedes any of this build work is squarely inside what we do. Our vCIO & IT Strategy Consulting service exists exactly for this kind of high-stakes infrastructure call: whether a data platform is actually justified yet, whether Fabric or an AWS-native build fits your team's existing skills and budget better, and what a realistic, phased build plan looks like before a single dollar goes toward a project with an 83% chance of running over budget or schedule. Once a direction is set, our Cloud Services & Migration team handles the infrastructure layer underneath any of these platforms -- network connectivity, identity and access management, backup and disaster recovery -- so the data platform decision doesn't happen in isolation from the rest of your IT environment. As a licensed California low-voltage contractor and managed service provider (License #1116987) with DIR public-works registration (PW-LR-1001158430), based in Carlsbad and serving San Diego County and businesses throughout California, we've seen enough infrastructure projects get sold before anyone checked whether they were actually needed. If your team is evaluating a data platform build and wants an outside read before committing budget, contact Mytek Pros at (619) 353-5702 or inquire@mytekpros.com.

Questions about it services? Get in touch or explore our IT Services.