# DMF Studio — Roadmap ## Next steps (in order) ### 1. WebSocket API in studio-manager Allow the graph to be changed at runtime without restarting. - Add a WebSocket server to `studio-manager` - API: load/reload graph, start/stop individual nodes, query status - `studio-manager` already has `load_graph()` — the WS layer calls it on demand and diffs against the running set (stop removed nodes, fork new ones) - **Required before**: frontend, live source switching, PiP (otherwise every graph change is a full restart) ### 2. Processing nodes — PiP / mixer First node that takes multiple input flows and produces an output flow. - Uses `mxlFlowSynchronizationGroup` to align grains from two inputs - Reference implementation: `nodes/testpattern` (writer) + `nodes/fakesink` (reader) - Only becomes useful with the WebSocket API (so you can switch sources live) ### 3. Vue.js frontend Visual graph editor that drives the WebSocket API. --- ## Redundancy Key constraint: **one writer per MXL flow** — can't run two identical nodes writing the same flow simultaneously. Redundancy lives at the pipeline level, not the node level. ### Dual pipeline on separate machines ``` Machine 1 (k8s node A) Machine 2 (k8s node B) decklinkin → [MXL] → ndiout decklinkin → [MXL] → ndiout ↓ ↓ (primary path) (backup path) \ / └──────→ [selector node] ←──────────┘ ↓ ndiout (final) ``` **Selector node** — reads two input flows, monitors grain validity flags, switches to backup when primary fails. Fits the existing node model; uses `mxlFlowSynchronizationGroup` to watch both flows. Key processing node to build once redundancy becomes a requirement. MXL shared memory requires all pods in a pipeline to be co-located on the same physical machine. Redundant pipelines naturally go on *different* machines — which is exactly right for hardware failure redundancy. ### What k8s gives for free - Stateless processing nodes (PiP, denoise, format convert): k8s restarts on crash, ~1-2 s gap — acceptable for non-critical path - `PodDisruptionBudget`: ensures critical nodes survive cluster maintenance - Leader election (k8s lease objects): two studio-managers, one active, one standby; automatic failover with no node code changes --- ## Kubernetes integration (mxl-k8s) Source: `~/codeproj/mxl-k8s` — not official, treat as reference, not truth. mxl-k8s is a full k8s control plane for MXL flows. Four runtime pieces: - **Operator** (Deployment): watches `MxlReceiver` CRDs, creates `MxlFlowMirror` per target node - **Agent** (DaemonSet): watches each node's MXL domain via `fanotify`, publishes `MxlFlow` CRDs with where flows live - **Gateway** (DaemonSet, `hostNetwork`): drives libmxl-fabrics RDMA/TCP between nodes — zero-copy grain transfer via registered mmap regions - **Shim** (`libmxl-intent.so`, LD_PRELOAD): intercepts `openat`/`stat`/`access` on `.mxl-flow/` paths in consumer pods; when a flow isn't local, asks the agent's UDS socket (`/run/mxl/agent.sock`) to materialize it via mirror, then retries — transparent to node code ### What changes for our nodes in k8s **Producer pods** (decklinkin, ndiin, testpattern): **zero code change**. - Add `hostPath: /run/mxl/domain` volume + `IPC_LOCK`, `SYS_RESOURCE` capabilities - `NODE_CONFIG` → Pod env var (from ConfigMap) - `MXL_DOMAIN` → `/run/mxl/domain` (standardized in k8s context) **Consumer pods, same node**: same as above, no code change. **Consumer pods, different node**: still no code change. - Add `initContainer` copying `libmxl-intent.so` from shim image - Set `LD_PRELOAD=/opt/mxl-intent/libmxl-intent.so` - Mount `/run/mxl` (whole dir, not just `/domain`) so agent socket is accessible - Create an `MxlReceiver` CRD pointing at the flow — operator handles the mirror plumbing ### What studio-manager becomes in k8s Currently: fork/exec child processes. In k8s: apply/delete Pods (or Deployments) with `NODE_CONFIG` env vars. For cross-node flows: create `MxlReceiver` CRDs instead of wiring flows manually. Same-node pipeline: all pods get `nodeAffinity: requiredDuringScheduling → same host`. Cross-node: add shim + `MxlReceiver`; mxl-k8s handles the rest. --- ## Architecture decisions ### Separate audio and video threads in nodes **Decision**: sink nodes (`decklinkout`, `ndiout`) and likely source nodes should process audio and video on separate threads. **Why**: audio and video have different timing granularities. - Video: one grain every ~40 ms (at 25 fps) — coarse, can block - Audio: must flow continuously at sample granularity — any stall causes dropout In the current single-thread model, video stalls (e.g. `TOO_EARLY` retries) pause audio too. In `decklinkout` this is especially bad — DeckLink's timestamped audio buffer underruns if it isn't fed consistently. **What it looks like**: - **Audio thread**: tight loop, continuously drains MXL audio ring buffer and pushes to output (DeckLink `ScheduleAudioSamples` / NDI send). No video logic. - **Video thread**: current main loop, handles grain read → process → output at frame rate. - **Shared state**: only `g_running` and the output handle. No frame data crosses the boundary — each thread reads its own MXL flow independently. MXL clock keeps them in sync without explicit A/V coordination. **When**: after the WebSocket API, when running real content and audio quality matters. Current single-thread model is acceptable for development.