Business

Databricks Releases Funke for Native HL7v2 Parsing

Databricks has launched Funke, an open-source Python and PySpark library that parses complex HL7v2 healthcare messages directly into native Spark formats on its Lakehouse platform.

Databricks AI16 hrs agoBusiness
Image: Databricks AI

Databricks has introduced Funke, an open-source tool designed to ingest and parse HL7v2 electronic health record messages directly within the Databricks Lakehouse. Funke serves as the modern successor to Smolder, a Scala-based library that Databricks released in 2021. While Smolder helped teams avoid hand-coding parsers, it predated modern platform features. Rebuilt in Python and PySpark, Funke integrates natively with Databricks' current architecture, including Unity Catalog for data governance, Declarative Automation Bundles for deployment, and Spark Declarative Pipelines for streaming ingestion.

The library addresses the long-standing difficulty of handling HL7v2 messages, which use highly nested, delimiter-encoded structures. Instead of forcing healthcare organizations to translate messages into FHIR formats or rely on third-party engines that flatten data into wide tables, Funke parses the messages directly into native Spark types. This lossless approach preserves the entire hierarchy of segments, fields, components, and subcomponents. Because the parsed data lands as a standard Spark column, developers and SQL analysts can query specific elements directly using positional index chains or built-in helper functions like get_value and get_segment.

Funke deploys as a Declarative Pipeline following the medallion architecture. Raw HL7 files arrive in a Unity Catalog volume, where Auto Loader ingests them into a bronze raw_messages table. The pipeline then parses this stream into a silver parsed_messages table, adding a structured hl7 column. To help users get started, the release includes a runnable demo featuring a synthetic admission, discharge, and transfer event generator. This demo showcases a complete pipeline that transforms raw clinical data into gold-level operational tables, such as a live hospital census and a bed utilization dashboard.

The tool is available as a Databricks Industry Solutions accelerator. It is open source under the Databricks License, allowing healthcare organizations to clone the repository and deploy it using the Databricks command-line interface or asset bundle editor.

This is our own summary of reporting by Databricks AI

More in Business