Integrations
Siemens S7 to Databricks, without a script.
Read the PLC's data blocks directly with MaestroHub's S7 connector, no OPC UA server needed in between, turn raw addresses into named values with units and quality, and write them to Databricks. If the link drops, writes wait on disk and are delivered when it returns.
The flow
One value, from the machine to Databricks. Illustrative.
Siemens S7PLC protocolDB10.DBD4
MaestroHubOven 3 temperature · 182.5 °C · Good
- named, unit, quality
- origin stamped
- buffered on disk
Databrickslakehouse/Volumes/plant/raw/oven3/Why do it with MaestroHub
More than a pipe from Siemens S7 to Databricks.
No gateway to buy and maintain
MaestroHub talks S7 to the PLC itself, so there is no separate OPC server or protocol gateway in the chain.
DB addresses become meaning
DB10.DBD4 becomes "Oven 3, temperature, °C" once, in one place, instead of in every notebook.
Built for the shop floor network
Runs next to the PLCs on a small edge box and keeps working through outages to the cloud.
What lands in Databricks
An example. You choose the fields in the pipeline.
| ts | asset | signal | value | unit | quality |
|---|---|---|---|---|---|
| 2026-10-03 06:00:01 | bake/oven3 | temperature | 182.5 | °C | good |
Set it up in three steps
- 1
Connect Siemens S7
Add the S7 connection with the PLC's address, rack and slot, then create read functions for the data blocks you need.
- 2
Give the values meaning
Map each value to a topic in the namespace with its name, unit and schema. Quality and origin are stamped on every value.
- 3
Deliver to Databricks
Add a Databricks output: write files to a Unity Catalog Volume, or insert rows through a SQL warehouse.
Before you start
What each side needs, from the connector documentation.
Siemens S7
- Network reach to the PLC on port 102 (ISO-on-TCP), entered as host:port.
- The CPU rack and slot from the PLC hardware configuration. Common combinations are rack 0 with slot 1 or slot 2.
- The session password, only if password protection is enabled in the PLC security settings. Maximum 8 characters.
- The DB numbers, byte offsets and data types of the values you need, for example from the TIA Portal project.
Example settings
- PLC Address
- 192.168.1.100:102
- Rack
- 0
- Slot
- 1
- Connection Type
- Programming Device (PG)
- Request Timeout (ms)
- 10000
- Idle Timeout (ms)
- 30000
Databricks
- A Databricks workspace URL starting with https://, for example https://myworkspace.cloud.databricks.com.
- A personal access token, or an OAuth M2M client ID and client secret.
- Unity Catalog enabled in the workspace, with the target volumes already created.
- For Databricks SQL: the SQL warehouse ID, found under Connection Details on the warehouse page.
Example settings
- Workspace URL
- https://<your-workspace>.cloud.databricks.com
- Auth Type
- OAuth M2M
- Client ID
- <your-client-id>
- Default Volume Path
- /Volumes/my_catalog/my_schema/my_volume/
- Max File Size (MB)
- 25
- Overwrite (Write function)
- false
Things to know
Pitfalls and limits the documentation calls out, so they don't surprise you on site.
Siemens S7
- Process Inputs (I) are read-only and writes to them are rejected. CHAR, STRING, WSTRING and the time and counter types can be written to DBs only, not to memory areas.
- For STRING and WSTRING, Size must be the declared maximum length plus the 2-byte S7 header. STRING[20] in TIA Portal is 22 bytes.
- A passing Test All checks the configuration only. Submit can still fail if the PLC rejects it, for example when the DB does not exist or the address is out of range on that PLC.
Limits
- Write functions handle one value each. Only reads group several data points into one function.
- TIMER, COUNTER, TIME_OF_DAY and DATE_AND_TIME can be read but not written.
Databricks
- Volume paths must follow /Volumes/<catalog>/<schema>/<volume>/. Set a Default Volume Path to use relative paths in functions.
- Write overwrites an existing file by default. Set Overwrite to false to make the call fail instead.
- In Databricks SQL, the connection's Query Timeout caps every statement. A function's Timeout can shorten it but not extend it.
Limits
- File reads and writes are limited by Max File Size: 25 MB by default, 124 MB at most.
- Databricks Storage works only on Unity Catalog Volumes.
What teams use it for
Process optimisation
Recipe parameters and outcomes per batch, analysed in Databricks.
Downtime analysis
Machine states and stop reasons from the PLC, aggregated per line and shift.
Digital twin data
A live, named feed of the line for simulation and what-if models.
Ways to connect Siemens S7 to Databricks
In general terms. Check any specific product for its own details.
Compare a custom script, a flow tool, a cloud vendor's edge service and MaestroHub
Swipe sideways to see every column
| A custom script | A flow tool | A cloud vendor's edge service | MaestroHub | |
|---|---|---|---|---|
| Talks Siemens S7 | A library you choose | Community plug-ins | Depends on the vendor | Built in, one of 90+ connectors |
| Names, units and schema | You write it | You build it | Partly, in the vendor's model | One governed namespace |
| Quality on every value | You write it | You build it | Depends on the vendor | Built in, carried through calculations |
| Where each value came from | You write it | You build it | Depends on the vendor | Stamped on every value |
| Survives a network outage | You build it | You build it | Usually, to that vendor's cloud | On disk, per destination, in order |
| Several destinations at once | One script each | Yes, flow by flow | Mostly that vendor's cloud | Any mix, each with its own buffer |
| Permissions and audit | You build it | You build it | Cloud account permissions | Roles, single sign-on, audit trail |
| AI agents can use the data | You build it | You build it | Depends on the vendor | Through the MCP server |
Questions
How do I connect Siemens S7 PLCs to Databricks?
Read the PLC's data blocks directly with MaestroHub's S7 connector, no OPC UA server needed in between, turn raw addresses into named values with units and quality, and write them to Databricks. If the link drops, writes wait on disk and are delivered when it returns.
Do I need to write code to connect Siemens S7 to Databricks?
No. You configure the Siemens S7 connection and the Databricks output in MaestroHub and join them with a pipeline. Transformations can be added where you need them.
Where does MaestroHub run for this?
On your own infrastructure next to the machines: an edge box, a virtual machine or Kubernetes. Only the rows you choose leave the plant for Databricks.
Why not just write a script for Siemens S7 to Databricks?
A script works on day one. The cost comes later: decoding and naming values, handling bad quality, buffering through outages, keeping credentials safe, and doing it again for the next machine and the next destination. MaestroHub does those once, for every connector.
What else can Siemens S7 data go to?
More than 90 connectors are included in every edition, among them historians, databases, cloud warehouses, message brokers and business systems, so the same values can feed several destinations at once.
Try Siemens S7 to Databricks yourself.
The free trial includes every connector. No factory to hand? The Digital Factory Simulator serves OPC UA, Modbus, MQTT and more on your laptop.


