News

Databricks releases Funke for native HL7v2 parsing

Databricks has released Funke, a Python and PySpark accelerator for parsing HL7v2 into native Spark structures within a Unity Catalog-governed ingestion pipeline.

D
Oct 10, 2026 · 2 min read

Databricks released Funke on October 8 as an open-source Python and PySpark accelerator that parses HL7v2 messages directly into native Spark structures. For healthcare data teams, that creates a path from raw clinical messages to queryable lakehouse data without first converting the messages into FHIR resources.

Funke succeeds Smolder, Databricks’ earlier Scala-based Spark SQL data source and helper library for loading and parsing HL7v2. Smolder exposed messages as DataFrames and included functions for extracting segments, fields and subfields. Funke brings that approach into current Databricks workflows with a Python and PySpark library and a deployable pipeline built around Databricks Asset Bundles, Spark Declarative Pipelines and Unity Catalog.

The Funke repository documents two ways to use the parser. Developers can parse an individual message in plain Python or apply a supplied PySpark user-defined function to a DataFrame column. The resulting Spark value is a map keyed by segment name. Repeated segments are retained, and fields remain addressable by field number, repetition, component and subcomponent. Databricks says the design is intended to preserve the message hierarchy as far into the pipeline as possible.

This direct representation is distinct from FHIR conversion. The FHIR specification defines FHIR as a resource-based standard for exchanging healthcare information. Funke does not create FHIR resources; it represents the original HL7v2 message hierarchy as Spark-native maps, arrays and structs. Teams that need FHIR still require a separate mapping and conversion step.

The packaged ingestion pattern creates two tables. New files land as plain text with ingestion metadata in raw_messages, while the parser adds a structured hl7 column in parsed_messages. Teams define downstream gold tables for their own use cases and can reach parsed elements through DataFrame or SQL expressions or Funke helper functions. The deployed schema, volume and tables sit under Unity Catalog governance.

Databricks describes Funke as a parser and ingestion accelerator, not an interface engine. The company also says the tool does not determine what a field means for a hospital’s business process. Clinical and HL7 expertise is still needed to map message fields to business concepts.

Databricks says Funke supports every HL7 message type and version and is designed for lossless parsing. The supplied materials, however, include neither an independent compatibility assessment nor losslessness testing across real-world customizations. They also provide no performance benchmark or adoption data. The repository makes Funke available under the Databricks License for exploration and says the project is offered as-is, without formal Databricks support or service-level agreements.

More news