Skip to content
MaestroHub
Menu
Free trialdownload and run it todayStart the pilot

Integrations

PI System to Databricks, without a script.

Keep the PI System as your historian and let MaestroHub read from it through the PI Web API: current, recorded, interpolated and summary values, browsed through the Asset Framework. MaestroHub then writes them to Databricks with names, units and quality.

The flow

One value, from the machine to Databricks. Illustrative.

PI Systemhistorian\\PI01\Plant\Line2|Temperature
MaestroHub

Line 2 temperature · 81.4 °C · Good

  • named, unit, quality
  • origin stamped
  • buffered on disk
Databrickslakehouse/Volumes/plant/pi/line2/

Why do it with MaestroHub

More than a pipe from PI System to Databricks.

01

PI stays the system of record

MaestroHub reads from PI rather than replacing it, so historian investments and AF models stay in use.

02

PI and non-PI data together

Machines and systems PI never collected join the same namespace before they reach Databricks.

03

No .NET side-car

The PI Web API connection is pure HTTPS, nothing extra to install on a Windows server.

What lands in Databricks

An example. You choose the fields in the pipeline.

tsassetsignalvalueunitquality
2026-10-03 06:00:01line2temperature81.4°Cgood

Set it up in three steps

  1. 1

    Connect PI System

    Add the PI Web API connection, browse the Asset Framework and pick the PI Points or attributes to read.

  2. 2

    Give the values meaning

    Map each value to a topic in the namespace with its name, unit and schema. Quality and origin are stamped on every value.

  3. 3

    Deliver to Databricks

    Add a Databricks output: write files to a Unity Catalog Volume, or insert rows through a SQL warehouse.

Before you start

What each side needs, from the connector documentation.

PI System

  • PI Web API 2017 or later, reachable from MaestroHub at its base URL, for example https://piserver.contoso.com/piwebapi.
  • An account with PI Web API permissions, using Anonymous, Basic, Bearer Token or Kerberos authentication.
  • For Kerberos: a keytab and krb5.conf readable on the MaestroHub host, network reach to the KDC, and a base URL FQDN that matches the service's HTTP/ SPN. No domain join is needed.
  • For Bearer Token: a JWT from PI Web API's configured trusted issuer, acquired and refreshed outside MaestroHub.

Example settings

Base URL
https://<your-pi-host>/piwebapi
Authentication Mode
Kerberos / SPNEGO
Keytab Path
/etc/krb5.keytab
Service Principal
HTTP/<your-pi-host>@<YOUR-REALM>
krb5.conf Path
/etc/krb5.conf
Verify TLS
true
PI System connector documentation

Databricks

  • A Databricks workspace URL starting with https://, for example https://myworkspace.cloud.databricks.com.
  • A personal access token, or an OAuth M2M client ID and client secret.
  • Unity Catalog enabled in the workspace, with the target volumes already created.
  • For Databricks SQL: the SQL warehouse ID, found under Connection Details on the warehouse page.

Example settings

Workspace URL
https://<your-workspace>.cloud.databricks.com
Auth Type
OAuth M2M
Client ID
<your-client-id>
Default Volume Path
/Volumes/my_catalog/my_schema/my_volume/
Max File Size (MB)
25
Overwrite (Write function)
false
Databricks connector documentation

Things to know

Pitfalls and limits the documentation calls out, so they don't surprise you on site.

PI System

  • HTTP 403 with valid credentials means missing permission or a wrong CSRF header. Change the CSRF header name on the Advanced tab only if your PI host expects a different one.
  • Subscribe polls PI Web API Stream Updates; it is not a server push. Events that happen entirely during a disconnection are not replayed. Add a periodic Get Recorded read for signals where gaps matter.
  • A type mismatch or a bad unit on a write comes back as HTTP 400 with PI's own message. Numeric and boolean strings are sent as real JSON types.

Limits

  • A recorded read returns at most about 150,000 events and a summary read at most 10,000 buckets. Split the time window for larger pulls.
  • Subscribe stream paths cannot use ((parameter)) templates; each subscription binds one fixed stream.

Databricks

  • Volume paths must follow /Volumes/<catalog>/<schema>/<volume>/. Set a Default Volume Path to use relative paths in functions.
  • Write overwrites an existing file by default. Set Overwrite to false to make the call fail instead.
  • In Databricks SQL, the connection's Query Timeout caps every statement. A function's Timeout can shorten it but not extend it.

Limits

  • File reads and writes are limited by Max File Size: 25 MB by default, 124 MB at most.
  • Databricks Storage works only on Unity Catalog Volumes.

What teams use it for

Historian data for data science

Years of PI history available to Databricks notebooks.

Asset Framework in the lakehouse

AF structure carried over as asset and signal names.

Hybrid reporting

PI values joined with ERP and quality data.

Ways to connect PI System to Databricks

In general terms. Check any specific product for its own details.

Compare a custom script, a flow tool, a cloud vendor's edge service and MaestroHub

Swipe sideways to see every column

A custom scriptA flow toolA cloud vendor's edge serviceMaestroHub
Talks PI SystemA library you chooseCommunity plug-insDepends on the vendorBuilt in, one of 90+ connectors
Names, units and schemaYou write itYou build itPartly, in the vendor's modelOne governed namespace
Quality on every valueYou write itYou build itDepends on the vendorBuilt in, carried through calculations
Where each value came fromYou write itYou build itDepends on the vendorStamped on every value
Survives a network outageYou build itYou build itUsually, to that vendor's cloudOn disk, per destination, in order
Several destinations at onceOne script eachYes, flow by flowMostly that vendor's cloudAny mix, each with its own buffer
Permissions and auditYou build itYou build itCloud account permissionsRoles, single sign-on, audit trail
AI agents can use the dataYou build itYou build itDepends on the vendorThrough the MCP server

Questions

How do I get data out of the PI System into Databricks?

Keep the PI System as your historian and let MaestroHub read from it through the PI Web API: current, recorded, interpolated and summary values, browsed through the Asset Framework. MaestroHub then writes them to Databricks with names, units and quality.

Do I need to write code to connect PI System to Databricks?

No. You configure the PI System connection and the Databricks output in MaestroHub and join them with a pipeline. Transformations can be added where you need them.

Where does MaestroHub run for this?

On your own infrastructure next to the machines: an edge box, a virtual machine or Kubernetes. Only the rows you choose leave the plant for Databricks.

Why not just write a script for PI System to Databricks?

A script works on day one. The cost comes later: decoding and naming values, handling bad quality, buffering through outages, keeping credentials safe, and doing it again for the next machine and the next destination. MaestroHub does those once, for every connector.

What else can PI System data go to?

More than 90 connectors are included in every edition, among them historians, databases, cloud warehouses, message brokers and business systems, so the same values can feed several destinations at once.

Try PI System to Databricks yourself.

The free trial includes every connector. No factory to hand? The Digital Factory Simulator serves OPC UA, Modbus, MQTT and more on your laptop.

Cookie preferences

Choose which categories of cookies you allow. Strictly necessary cookies are always on because the site cannot function without them.