Integrations
PI System to Databricks, without a script.
Keep the PI System as your historian and let MaestroHub read from it through the PI Web API: current, recorded, interpolated and summary values, browsed through the Asset Framework. MaestroHub then writes them to Databricks with names, units and quality.
The flow
One value, from the machine to Databricks. Illustrative.
PI Systemhistorian\\PI01\Plant\Line2|Temperature
MaestroHubLine 2 temperature · 81.4 °C · Good
- named, unit, quality
- origin stamped
- buffered on disk
Databrickslakehouse/Volumes/plant/pi/line2/Why do it with MaestroHub
More than a pipe from PI System to Databricks.
PI stays the system of record
MaestroHub reads from PI rather than replacing it, so historian investments and AF models stay in use.
PI and non-PI data together
Machines and systems PI never collected join the same namespace before they reach Databricks.
No .NET side-car
The PI Web API connection is pure HTTPS, nothing extra to install on a Windows server.
What lands in Databricks
An example. You choose the fields in the pipeline.
| ts | asset | signal | value | unit | quality |
|---|---|---|---|---|---|
| 2026-10-03 06:00:01 | line2 | temperature | 81.4 | °C | good |
Set it up in three steps
- 1
Connect PI System
Add the PI Web API connection, browse the Asset Framework and pick the PI Points or attributes to read.
- 2
Give the values meaning
Map each value to a topic in the namespace with its name, unit and schema. Quality and origin are stamped on every value.
- 3
Deliver to Databricks
Add a Databricks output: write files to a Unity Catalog Volume, or insert rows through a SQL warehouse.
Before you start
What each side needs, from the connector documentation.
PI System
- PI Web API 2017 or later, reachable from MaestroHub at its base URL, for example https://piserver.contoso.com/piwebapi.
- An account with PI Web API permissions, using Anonymous, Basic, Bearer Token or Kerberos authentication.
- For Kerberos: a keytab and krb5.conf readable on the MaestroHub host, network reach to the KDC, and a base URL FQDN that matches the service's HTTP/ SPN. No domain join is needed.
- For Bearer Token: a JWT from PI Web API's configured trusted issuer, acquired and refreshed outside MaestroHub.
Example settings
- Base URL
- https://<your-pi-host>/piwebapi
- Authentication Mode
- Kerberos / SPNEGO
- Keytab Path
- /etc/krb5.keytab
- Service Principal
- HTTP/<your-pi-host>@<YOUR-REALM>
- krb5.conf Path
- /etc/krb5.conf
- Verify TLS
- true
Databricks
- A Databricks workspace URL starting with https://, for example https://myworkspace.cloud.databricks.com.
- A personal access token, or an OAuth M2M client ID and client secret.
- Unity Catalog enabled in the workspace, with the target volumes already created.
- For Databricks SQL: the SQL warehouse ID, found under Connection Details on the warehouse page.
Example settings
- Workspace URL
- https://<your-workspace>.cloud.databricks.com
- Auth Type
- OAuth M2M
- Client ID
- <your-client-id>
- Default Volume Path
- /Volumes/my_catalog/my_schema/my_volume/
- Max File Size (MB)
- 25
- Overwrite (Write function)
- false
Things to know
Pitfalls and limits the documentation calls out, so they don't surprise you on site.
PI System
- HTTP 403 with valid credentials means missing permission or a wrong CSRF header. Change the CSRF header name on the Advanced tab only if your PI host expects a different one.
- Subscribe polls PI Web API Stream Updates; it is not a server push. Events that happen entirely during a disconnection are not replayed. Add a periodic Get Recorded read for signals where gaps matter.
- A type mismatch or a bad unit on a write comes back as HTTP 400 with PI's own message. Numeric and boolean strings are sent as real JSON types.
Limits
- A recorded read returns at most about 150,000 events and a summary read at most 10,000 buckets. Split the time window for larger pulls.
- Subscribe stream paths cannot use ((parameter)) templates; each subscription binds one fixed stream.
Databricks
- Volume paths must follow /Volumes/<catalog>/<schema>/<volume>/. Set a Default Volume Path to use relative paths in functions.
- Write overwrites an existing file by default. Set Overwrite to false to make the call fail instead.
- In Databricks SQL, the connection's Query Timeout caps every statement. A function's Timeout can shorten it but not extend it.
Limits
- File reads and writes are limited by Max File Size: 25 MB by default, 124 MB at most.
- Databricks Storage works only on Unity Catalog Volumes.
What teams use it for
Historian data for data science
Years of PI history available to Databricks notebooks.
Asset Framework in the lakehouse
AF structure carried over as asset and signal names.
Hybrid reporting
PI values joined with ERP and quality data.
Ways to connect PI System to Databricks
In general terms. Check any specific product for its own details.
Compare a custom script, a flow tool, a cloud vendor's edge service and MaestroHub
Swipe sideways to see every column
| A custom script | A flow tool | A cloud vendor's edge service | MaestroHub | |
|---|---|---|---|---|
| Talks PI System | A library you choose | Community plug-ins | Depends on the vendor | Built in, one of 90+ connectors |
| Names, units and schema | You write it | You build it | Partly, in the vendor's model | One governed namespace |
| Quality on every value | You write it | You build it | Depends on the vendor | Built in, carried through calculations |
| Where each value came from | You write it | You build it | Depends on the vendor | Stamped on every value |
| Survives a network outage | You build it | You build it | Usually, to that vendor's cloud | On disk, per destination, in order |
| Several destinations at once | One script each | Yes, flow by flow | Mostly that vendor's cloud | Any mix, each with its own buffer |
| Permissions and audit | You build it | You build it | Cloud account permissions | Roles, single sign-on, audit trail |
| AI agents can use the data | You build it | You build it | Depends on the vendor | Through the MCP server |
Questions
How do I get data out of the PI System into Databricks?
Keep the PI System as your historian and let MaestroHub read from it through the PI Web API: current, recorded, interpolated and summary values, browsed through the Asset Framework. MaestroHub then writes them to Databricks with names, units and quality.
Do I need to write code to connect PI System to Databricks?
No. You configure the PI System connection and the Databricks output in MaestroHub and join them with a pipeline. Transformations can be added where you need them.
Where does MaestroHub run for this?
On your own infrastructure next to the machines: an edge box, a virtual machine or Kubernetes. Only the rows you choose leave the plant for Databricks.
Why not just write a script for PI System to Databricks?
A script works on day one. The cost comes later: decoding and naming values, handling bad quality, buffering through outages, keeping credentials safe, and doing it again for the next machine and the next destination. MaestroHub does those once, for every connector.
What else can PI System data go to?
More than 90 connectors are included in every edition, among them historians, databases, cloud warehouses, message brokers and business systems, so the same values can feed several destinations at once.
Try PI System to Databricks yourself.
The free trial includes every connector. No factory to hand? The Digital Factory Simulator serves OPC UA, Modbus, MQTT and more on your laptop.


