To start the backend run (add the --build flag to build images before starting containers (build from scratch)):
docker-compose up
docker-compose up --buildThis will build the GraphDB docker image and the FastAPI docker image.
The GraphDB instance will be available at: localhost:7200
The FastAPI documentation will be exposed at: http://0.0.0.0:9017
Desription of how to start the backend locally outside docker. The backend consists of a GraphDB database and a FastAPI server.
To start GraphDB, run these scripts from the GraphDB folder:
bash graphdb_install.sh
bash graphdb_preload.sh
bash graphdb_start.shTo test your graphbd connection run from your base folder (/sindit):
python run_test.pyGo to localhost:7200 to configure graphdb
First, ensure dependencies are installed:
uv pip install -e .or
uv syncTo start the FastAPI server, run:
python run_sindit.py{
"version": "0.2.0",
"configurations": [
{
"name": "Python Debugger: Current File",
"type": "debugpy",
"request": "launch",
"program": "${file}",
"console": "integratedTerminal",
"cwd": "${workspaceFolder}/src/sindit",
"env": {
"PYTHONPATH": "${workspaceFolder}/src"
},
"justMyCode": false
}
]
}The API requires a valid authentication token for most endpoints. Follow these steps to authenticate and use the API:
-
Generate a Token:
- Use the
/tokenendpoint to generate an access token. - Example
curlcommand:curl -X POST "http://127.0.0.1:9017/token" \ -H "Content-Type: application/x-www-form-urlencoded" \ -d "username=new_user&password=new_password"
- Replace
new_userandnew_passwordwith the credentials provided below.
- Use the
-
Use the Token:
- Include the token in the
Authorizationheader for all subsequent API calls:curl -X GET "http://127.0.0.1:9017/endpoint" \ -H "Authorization: Bearer your_generated_token_here"
- Include the token in the
-
Access API Documentation:
- The FastAPI documentation is available at:
http://127.0.0.1:9017/docs
- The FastAPI documentation is available at:
To add a new user, update the fake_users_db in authentication_endpoints.py with the following credentials:
fake_users_db = {
"new_user": {
"username": "new_user",
"full_name": "New User",
"email": "new_user@example.com",
"hashed_password": "$2b$12$eW5j9GdY3.EciS3oKQxJjOyIpoUNiFZxrON4SXt3wVrgSbE1gDMba", # Password: new_password
"disabled": False,
}
}To generate a new hashed password, use the Python snippet in password_hash.py.
Replace "new_password" with your desired password.
This section documents the AI layer added on top of SINDIT for industrial CNC machine analysis. It is composed of a multi-agent architecture that answers natural-language questions by routing them to the right data source.
User question (natural language)
│
▼
IntegratedAgent ←── semantic classification (embeddings + cosine similarity)
│
├─── RetrieverAgent → Knowledge Graph (Text2SPARQL → GraphDB)
├─── AnalyticsAgent → Historical CNC data (Parquet files)
├─── MonitoringAgent → Real-time MQTT data (SINTEF broker)
└─── GreenAgent → Carbon footprint (DPP files / REST API)
The IntegratedAgent classifies each question and delegates to one or several specialised agents:
| Agent | Data source | Use case |
|---|---|---|
| RetrieverAgent | GraphDB knowledge graph | Documentation, machine specs, component descriptions |
| AnalyticsAgent | Parquet files (data/cnc/) |
Historical analysis of machined workpieces |
| MonitoringAgent | MQTT broker (real-time) | Live sensor values, current machine state |
| GreenAgent | DPP JSON files / REST API | Carbon footprint per production batch (ISO 14044, phases A1/A2/A3) |
The LLM (Ollama) runs on the SINTEF mainframe. To expose it locally, open an SSH tunnel before starting the chatbot:
ssh -L 11434:localhost:11434 example@mainframe.sintef.noThis forwards localhost:11434 on your machine to the Ollama server on the mainframe. The tunnel must stay open for the entire session.
streamlit run ui/src/chatbot/chatbot_ui.py --server.port 8502Accessible at: http://localhost:8502
Example questions:
What is the description of the "Pneumatic System"?
What did I ask you in my previous question?
Analyse the workpiece OF10001 from 06:00 to 06:30
Show vibration for OF10005 from 2026-03-13 10:18 to 14:00
Analyse OF10001 from 06:00 to 08:00
What is the carbon footprint of OF10005?
Compare the carbon footprint of OF10004 and OF10005.
Which batch has the highest CO2 emissions?
Standalone Streamlit app for interactive exploration of workpiece data (time series, tool changes, temperatures, chatter statistics):
streamlit run sindit/src/sindit/processing/viz.pyTechnical PDF documents (machine manuals, spare parts lists, etc.) can be ingested and stored in the knowledge graph. The LLM extracts entities (assets, properties, relationships) from each text chunk and stores them via the SINDIT REST API.
Before running, make sure the SSH tunnel is open (Ollama needed) and the SINDIT API is running (localhost:9017).
python -m sindit.src.sindit.knowledge_graph.pdf_file_to_kgConfiguration is done at the top of the __main__ block in pdf_file_to_kg.py:
PDF_PATH = "data/documents/my_machine.pdf"
START_PAGE = 5 # skip first N pages (cover, table of contents…)
START_CHUNK = 0 # resume from chunk N after a crash (0 = start fresh)
CLEAN_GRAPH = False # True = wipe the KG before starting
GRAPH_NAME = "default"The GreenAgent answers natural-language questions about the environmental impact of CNC production batches. It reads Digital Product Passport (DPP) files, structured JSON documents following the ISO 14044 standard, and uses an LLM to generate a precise, context-aware response. The agent covers four lifecycle phases:
| Phase | Description |
|---|---|
| A1 | Raw material extraction |
| A2 | Transport to the factory |
| A3 | Manufacturing (energy consumed by the CNC machine) |
| A1–A3 | Total cradle-to-gate footprint |
User question (e.g. "Compare the carbon footprint of OF10004 and OF10005")
│
▼
1. Batch ID extraction
└── The agent scans the question for OF-series identifiers (e.g. OF10004, OF10005)
│
▼
2. DPP data loading
├── Mock mode (USE_MOCK=True) → reads local JSON files under data/cnc/<batch>/
└── Live mode (USE_MOCK=False) → fetches data from the REST API at DPP_API_URL
│
▼
3. Single batch or multi-batch?
├── Single batch → generates a plain-text summary, passes it to the LLM
└── Multiple batches → builds a comparative context (totals, dominant phases,
highest/lowest emitter) AND generates a grouped bar chart (Plotly)
│
▼
4. LLM answer generation
├── Single batch → QA_PROMPT (ISO 14044, units in kg CO2eq)
└── Multi-batch → MULTI_BATCH_QA_PROMPT + LLMImageAnalyzer (if chart available)
│
▼
5. Response returned to IntegratedAgent
└── self.last_figure is set so chatbot_ui can render the interactive chart
Each production batch that has environmental data contains one DPP file:
data/cnc/OF10004/DPP_goimek_OF10004_G_BQC_S8CF2G_2026-06-01.json
data/cnc/OF10005/DPP_goimek_OF10005_G_BQC_S8CF2G_2026-06-01.json
The relevant structure inside each file is:
{
"DPP": [{
"company": "Goimek",
"product": "...",
"carbonFootprint": {
"ProductCarbonFootprints": [
{ "LifeCyclePhases": [{"LifeCyclePhase": "A1"}], "PCF in kg CO2eq": 16819.5 },
{ "LifeCyclePhases": [{"LifeCyclePhase": "A2"}], "PCF in kg CO2eq": 227.7 },
{ "LifeCyclePhases": [{"LifeCyclePhase": "A3"}], "PCF in kg CO2eq": 921.0 },
{ "LifeCyclePhases": [{"LifeCyclePhase": "A1-A3"}], "PCF in kg CO2eq": 17968.3 }
]
}
}]
}To switch from local JSON files to a live API, set USE_MOCK=False in the environment and configure DPP_API_URL.
When the question targets multiple batches, the GreenAgent does not limit its output to a text answer. It also produces a grouped bar chart that makes phase-level differences between batches immediately visible.
The chart pipeline works as follows:
-
Data extraction —
_parse_footprints(data)iterates over theProductCarbonFootprintsarray of each DPP and builds a{phase: float}dictionary (e.g.{"A1": 16819.5, "A2": 227.7, "A3": 921.0}). Only individual phases (A1, A2, A3) are plotted; the aggregate A1–A3 total is reported in the text instead. -
Figure construction —
_generate_co2_chart()uses Plotly (plotly.graph_objects) to build a grouped bar chart. Each production batch maps to onego.Bartrace; the three lifecycle phases are placed on the x-axis. Value labels are displayed above each bar for direct readability. -
PNG export —
save_chart_as_image()(imported fromllm_image_explanation) converts the Plotly figure to a PNG file stored undergenerated_charts/. The filename includes a timestamp to avoid overwriting previous charts. -
Visual analysis by LLM —
LLMImageAnalyzer.analyze_with_data_image_and_custom_prompt()sends the PNG to a vision-capable LLM instance together with the structured text context and theMULTI_BATCH_QA_PROMPT. This means the LLM's answer is grounded both in the raw numerical data (text context) and in the visual representation (image), which improves the quality of comparative observations. -
Fallback — if
save_chart_as_image()returnsFalse(e.g., a display server is unavailable), the agent falls back to a text-only response usingMULTI_BATCH_QA_PROMPTwithout the image. The LLM answer is still generated; only the chart is absent.
_generate_co2_chart()
│
├── _parse_footprints() → {A1: x, A2: y, A3: z} per batch
├── go.Figure + go.Bar → grouped Plotly figure
├── save_chart_as_image() → generated_charts/<name>_<ts>.png
│ │
│ └── success? ──► LLMImageAnalyzer.analyze_with_data_image_and_custom_prompt()
│ (text context + PNG → LLM vision answer)
└── failure? ──────────► MULTI_BATCH_QA_PROMPT text-only fallback
The GreenAgent does not display results directly — it is always called by the IntegratedAgent, which orchestrates the response and manages the UI layer.
To allow the chart to reach the Streamlit chatbot interface, GreenAgent exposes a shared state attribute:
self.last_figure = fig # set inside _generate_co2_chart()After GreenAgent._run() returns, the IntegratedAgent reads self.green_agent.last_figure. If a Plotly figure is present, it passes it up to chatbot_ui.py, which calls st.plotly_chart() to render the interactive chart alongside the text answer.
This design keeps the GreenAgent stateless with respect to the UI: it does not import Streamlit or any display dependency. The figure is produced as a pure Plotly object; display is entirely the responsibility of the caller. The pattern allows the same agent to be reused in non-Streamlit contexts (REST endpoint, CLI) without modification.
GreenAgent._run()
└── sets self.last_figure = fig
│
▼
IntegratedAgent reads last_figure
│
▼
chatbot_ui.py → st.plotly_chart(fig) ← interactive chart in browser
| Variable | Default | Description |
|---|---|---|
USE_MOCK |
True |
If True, reads local DPP JSON files. If False, calls the REST API. |
DPP_API_URL |
http://localhost:8888 |
Base URL of the live DPP REST API (only used when USE_MOCK=False). |
What is the carbon footprint of OF10005?
Compare the carbon footprint of OF10004 and OF10005.
Which batch has the highest CO2 emissions?
What percentage of OF10004's footprint comes from manufacturing?
CNC data is stored under data/cnc/ with one subfolder per work order (OF):
data/cnc/
├── OF10001/
│ ├── OF10001_G_BQC_S8CF2G_TYZBPS.csv ← general machine state
│ ├── OF10001_G_BQC_S8CF2G_TYZBPS.parquet
│ ├── OF10001_G_BQC_S8CF2G_BXCZ3M.csv ← axis power
│ ├── OF10001_G_BQC_S8CF2G_BXCZ3M.parquet
│ ├── OF10001_G_BQC_S8CF2G_7N4ZJ8.csv ← vibration & chatter
│ └── OF10001_G_BQC_S8CF2G_7N4ZJ8.parquet
├── OF10002/ …
├── OF10004/
│ ├── …
│ └── DPP_goimek_OF10004_G_BQC_S8CF2G_2026-06-01.json ← carbon footprint (DPP)
├── OF10005/
│ ├── …
│ └── DPP_goimek_OF10005_G_BQC_S8CF2G_2026-06-01.json ← carbon footprint (DPP)
DPP files follow the ISO 14044 structure with four lifecycle phases: A1 (raw material extraction), A2 (transport), A3 (manufacturing), and A1-A3 (total cradle-to-gate). They are consumed by the GreenAgent and can be replaced by a live REST API endpoint by setting USE_MOCK=False in the environment.
Naming convention: OFxxxxx_MachineID_SensorType.csv
Join key across files: timestamp (ISO 8601 UTC)
| OF | Start | End | Duration |
|---|---|---|---|
| OF10001 | 2025-09-01 06:15 | 2025-09-02 00:22 | ~18 h |
| OF10002 | 2025-09-02 01:10 | 2025-09-02 16:17 | ~15 h |
| OF10003 | 2025-09-02 16:58 | 2025-09-03 08:41 | ~16 h |
| OF10004 | 2025-10-22 00:19 | 2025-10-22 13:23 | ~13 h |
| OF10005 | 2026-03-13 10:18 | 2026-03-16 00:54 | ~62 h |
Raw CSV files must be converted to Parquet before the agents can use them (PyArrow predicate pushdown requires Parquet):
python -c "from sindit.src.sindit.processing.cnc_preprocessing import convert_all_cnc_data; convert_all_cnc_data()"This iterates over all data/cnc/OFxxxxx/*_TYZBPS.csv, *_BXCZ3M.csv, and *_7N4ZJ8.csv files and writes the corresponding .parquet alongside each CSV.
This is the main file. It captures the complete machine state at each timestamp: what it is doing, where it is, at what speed, with which tool, and how much energy it consumes.
| Column | Description |
|---|---|
timestamp |
Timestamp — join key across the 3 files |
Spindle_Speed_Actual |
Actual spindle speed (rpm) — > 0 means the machine is actively cutting |
Spindle_Speed_Commanded |
Spindle speed commanded by the NC program |
Spindle_Speed_Override |
Override applied by the operator (%) |
Feed_Rate_Actual |
Actual table feed rate (mm/min) — indicates a cutting move |
Feed_Rate_Commanded |
Feed rate commanded by the NC program |
Feed_Override |
Override applied by the operator (%) |
Position_MCS_X/Y/Z |
Absolute position in the Machine Coordinate System on the 3 linear axes |
Position_MCS_A/C |
Position of the rotary axes (5-axis machines only) |
Offset_X/Y/Z |
Axis offsets |
Power_Active |
Total active power consumed (W) |
Power_Apparent |
Apparent power (VA) |
Power_Reactive |
Reactive power (VAR) |
Power_Factor |
Power factor |
Power_Spindle |
Power consumed by the spindle motor alone (W) |
Energy_Total |
Cumulative energy counter (kWh) |
Tool_Number |
Active tool number (identifies which cutter is mounted) |
Tool_Length |
Active tool length (mm) |
Tool_Radius |
Active tool radius (mm) |
Program_Name |
Name of the NC program currently running |
Program_Block_Number |
Current NC block number (N…) |
Temperature_Head |
Head temperature (°C) |
Temperature_Room |
Room temperature (°C) |
Temperature_Y |
Y-axis temperature (°C) |
Temperature_Z |
Z-axis temperature (°C) |
Head_Angular_On |
Flag: angular head mode active |
Head_Auto_On |
Flag: automatic mode active |
Head_Boring_On |
Flag: boring mode active |
Operation_Mode |
Machine mode (automatic, manual, MDI…) |
Operation_Status |
Operation status (running, paused, emergency stop…) |
MCS = Machine Coordinate System — the absolute reference frame fixed to the machine itself, as opposed to the workpiece coordinate system.
Spindle — the rotary motor that spins the cutting tool at high speed (1 000 – 20 000 rpm). WhenSpindle_Speed_Actual > 0, the machine is actively cutting.
Lighter file that provides per-axis power consumption and the machine's operational mode. Useful for identifying which axis is under load.
| Column | Description |
|---|---|
timestamp |
Timestamp — join key across the 3 files |
Power_X1 |
Power consumed by motor 1 of the X axis (W) |
Power_X2 |
Power consumed by motor 2 of the X axis (W) |
Power_Y |
Power consumed by the Y axis (W) |
Power_Z |
Power consumed by the Z axis (W) |
Operation_Mode |
Machine mode (automatic, manual, MDI…) |
Operation_Status |
Operation status (running, paused, emergency stop…) |
Two X motors: the X axis uses a dual-drive configuration (one motor on each side of the gantry) to prevent racking over long travels. Each motor is monitored independently.
The most detailed file. It contains the full vibration signature of the machine: severity index, harmonic decomposition (FFT), spectral peaks, and a binary chatter detection flag.
| Column | Description |
|---|---|
timestamp |
Timestamp — join key across the 3 files |
Chatter_Detection_OnOff_X/Y |
Binary flag: chatter detected on axis X or Y |
Chatter_Detection_Amplitude_X/Y |
Amplitude of the detected chatter signal |
Chatter_Detection_Frequency_X/Y |
Frequency of the detected chatter (Hz) |
Vibration_Harmonic_N_X/Y_Amplitude |
Amplitude of the N-th harmonic (N = 1 to 8) |
Vibration_Harmonic_N_X/Y_Frequency |
Frequency of the N-th harmonic (Hz) |
Vibration_Peak_N_X/Y_Amplitude |
Amplitude of the N-th spectral peak (N = 1 to 5) |
Vibration_Peak_N_X/Y_Frequency |
Frequency of the N-th spectral peak (Hz) |
Vibration_Severity_X/Y |
Global vibration severity index (summary of all peaks) |
Chatter is a self-sustained resonant vibration between the tool and the workpiece. It degrades surface quality and accelerates tool wear. Chatter_Detection_OnOff_X/Y = True is an immediate alert.
Harmonics (32 columns, Harmonic_1 to Harmonic_8): the FFT of the vibration signal is decomposed into its fundamental frequency and its integer multiples. Harmonic_1 typically corresponds to the spindle rotation frequency; higher harmonics indicate more complex dynamics.
Peaks (20 columns, Peak_1 to Peak_5): the 5 frequencies with the highest amplitude in the full spectrum, regardless of their harmonic relationship. An unexpected peak at an unusual frequency may indicate a worn tool, loose fixture, or failing bearing.
Vibration_Severity: a single scalar summarising overall vibration intensity (computed following industrial norms such as ISO 10816). A low value means normal machining; a high value means excessive vibration → risk of poor surface finish or tool breakage.
