NVIDIA’s open-source samples show how exported models can run with hardware acceleration on Windows and Linux.
NVIDIA describes DIN Deploy as an open-source collection of practical C++ samples for developers building local AI applications. The project connects model export, ONNX Runtime, and NVIDIA’s TensorRT RTX execution provider in one workflow.
Each sample begins with a Python exporter. It downloads a model checkpoint from Hugging Face and converts it into an ONNX artifact, a portable model file that can move between compatible tools. A native C++ command-line application then loads that artifact through ONNX Runtime, the software layer that runs the model.
That split keeps model conversion separate from application code. Most shared code uses ONNX Runtime’s session and tensor interfaces, while vendor-specific CUDA code and kernels appear only in optional accelerated paths. In practice, developers can start with a common C++ interface and add acceleration where supported.
The repository includes CMake presets for Windows and Linux, including Arm64 variants. DirectX is available only on Windows. A FLUX.2 image-generation sample also uses Vulkan and DirectX to connect graphics resources with model processing.
DIN Deploy includes offline and streaming speech recognition, interactive image and video segmentation, and prompt-driven image generation. NVIDIA’s table reports that, on DGX Spark, Parakeet speech recognition ran at 206.41 times real time on the GPU versus 14.44 times on the CPU. A SAM 2.1 example reached 38.3 frames per second on the GPU versus 0.5 on the CPU.
Those results describe NVIDIA’s measurements on DGX Spark. Developers evaluating the samples should build them with the provided presets, then measure performance on their own target hardware.
