Hospital for Special Surgery (HSS), the New York orthopedic academic medical center, consolidated 40-plus source systems and more than 14,500 tables across 130-plus schemas into a single Databricks lakehouse over a 10-month programme, giving finance, operations, clinical research, and quality teams one trusted view of the business while preparing the data foundation for AI-driven analytics.
The lakehouse was built because HSS had been running on a fragmented analytics stack — multiple departmental warehouses, separate clinical and financial data marts, and a proliferation of one-off data pipelines — that left every team reconciling against slightly different numbers and made cross-functional questions about cost, capacity, and patient outcomes impossible to answer reliably. Reporting cycles on quarterly finance-and-operations reviews routinely took weeks because analysts had to stitch together extracts from several systems. The data platform team selected Databricks because the lakehouse pattern allows HSS to unify structured warehouse tables, semi-structured clinical notes, and unstructured imaging metadata in one governed layer, with Unity Catalog providing the cross-domain access controls that a HIPAA-regulated academic medical center requires.
The technical architecture uses Databricks Delta Live Tables to ingest from the source systems via Fivetran and HSS's own CDC pipelines, applies dbt for SQL transformations, and serves curated datasets to Tableau for finance and operations, to custom internal applications for clinical research, and to a growing set of Python notebooks used by the data science team. Unity Catalog governs access so that protected health information is masked for most analytics users and only unblocked for credentialed clinical-research workloads. The migration was staged by domain — finance first, then operations, then clinical research — and the team retired three legacy warehouses during the programme.
The operational impact is that finance close cycles have compressed because analysts no longer reconcile against separate systems, and operational leaders can answer cross-domain questions about cost-per-episode, OR utilisation, and length-of-stay in a single query. The data platform team has shifted time previously spent on pipeline maintenance to building AI-driven use cases on top of the lakehouse, with initial priorities in clinical operations and quality measurement. HSS plans to extend the platform with additional machine-learning workloads, including predictive models for readmission risk and OR scheduling, now that a single governed data layer is in place.