5.6 KiB
DMF Studio — Roadmap
Next steps (in order)
1. WebSocket API in studio-manager
Allow the graph to be changed at runtime without restarting.
- Add a WebSocket server to
studio-manager - API: load/reload graph, start/stop individual nodes, query status
studio-manageralready hasload_graph()— the WS layer calls it on demand and diffs against the running set (stop removed nodes, fork new ones)- Required before: frontend, live source switching, PiP (otherwise every graph change is a full restart)
2. Processing nodes — PiP / mixer
First node that takes multiple input flows and produces an output flow.
- Uses
mxlFlowSynchronizationGroupto align grains from two inputs - Reference implementation:
nodes/testpattern(writer) +nodes/fakesink(reader) - Only becomes useful with the WebSocket API (so you can switch sources live)
3. Vue.js frontend
Visual graph editor that drives the WebSocket API.
Redundancy
Key constraint: one writer per MXL flow — can't run two identical nodes writing the same flow simultaneously. Redundancy lives at the pipeline level, not the node level.
Dual pipeline on separate machines
Machine 1 (k8s node A) Machine 2 (k8s node B)
decklinkin → [MXL] → ndiout decklinkin → [MXL] → ndiout
↓ ↓
(primary path) (backup path)
\ /
└──────→ [selector node] ←──────────┘
↓
ndiout (final)
Selector node — reads two input flows, monitors grain validity flags, switches
to backup when primary fails. Fits the existing node model; uses
mxlFlowSynchronizationGroup to watch both flows. Key processing node to build
once redundancy becomes a requirement.
MXL shared memory requires all pods in a pipeline to be co-located on the same physical machine. Redundant pipelines naturally go on different machines — which is exactly right for hardware failure redundancy.
What k8s gives for free
- Stateless processing nodes (PiP, denoise, format convert): k8s restarts on crash, ~1-2 s gap — acceptable for non-critical path
PodDisruptionBudget: ensures critical nodes survive cluster maintenance- Leader election (k8s lease objects): two studio-managers, one active, one standby; automatic failover with no node code changes
Kubernetes integration (mxl-k8s)
Source: ~/codeproj/mxl-k8s — not official, treat as reference, not truth.
mxl-k8s is a full k8s control plane for MXL flows. Four runtime pieces:
- Operator (Deployment): watches
MxlReceiverCRDs, createsMxlFlowMirrorper target node - Agent (DaemonSet): watches each node's MXL domain via
fanotify, publishesMxlFlowCRDs with where flows live - Gateway (DaemonSet,
hostNetwork): drives libmxl-fabrics RDMA/TCP between nodes — zero-copy grain transfer via registered mmap regions - Shim (
libmxl-intent.so, LD_PRELOAD): interceptsopenat/stat/accesson.mxl-flow/paths in consumer pods; when a flow isn't local, asks the agent's UDS socket (/run/mxl/agent.sock) to materialize it via mirror, then retries — transparent to node code
What changes for our nodes in k8s
Producer pods (decklinkin, ndiin, testpattern): zero code change.
- Add
hostPath: /run/mxl/domainvolume +IPC_LOCK,SYS_RESOURCEcapabilities NODE_CONFIG→ Pod env var (from ConfigMap)MXL_DOMAIN→/run/mxl/domain(standardized in k8s context)
Consumer pods, same node: same as above, no code change.
Consumer pods, different node: still no code change.
- Add
initContainercopyinglibmxl-intent.sofrom shim image - Set
LD_PRELOAD=/opt/mxl-intent/libmxl-intent.so - Mount
/run/mxl(whole dir, not just/domain) so agent socket is accessible - Create an
MxlReceiverCRD pointing at the flow — operator handles the mirror plumbing
What studio-manager becomes in k8s
Currently: fork/exec child processes. In k8s: apply/delete Pods (or Deployments) with NODE_CONFIG env vars. For cross-node flows: create MxlReceiver CRDs instead of wiring flows manually.
Same-node pipeline: all pods get nodeAffinity: requiredDuringScheduling → same host.
Cross-node: add shim + MxlReceiver; mxl-k8s handles the rest.
Architecture decisions
Separate audio and video threads in nodes
Decision: sink nodes (decklinkout, ndiout) and likely source nodes should
process audio and video on separate threads.
Why: audio and video have different timing granularities.
- Video: one grain every ~40 ms (at 25 fps) — coarse, can block
- Audio: must flow continuously at sample granularity — any stall causes dropout
In the current single-thread model, video stalls (e.g. TOO_EARLY retries)
pause audio too. In decklinkout this is especially bad — DeckLink's timestamped
audio buffer underruns if it isn't fed consistently.
What it looks like:
- Audio thread: tight loop, continuously drains MXL audio ring buffer and
pushes to output (DeckLink
ScheduleAudioSamples/ NDI send). No video logic. - Video thread: current main loop, handles grain read → process → output at frame rate.
- Shared state: only
g_runningand the output handle. No frame data crosses the boundary — each thread reads its own MXL flow independently. MXL clock keeps them in sync without explicit A/V coordination.
When: after the WebSocket API, when running real content and audio quality matters. Current single-thread model is acceptable for development.