ACMX 2.139.0
Dual-Backend Real-Time GPU Video Synthesis
Loading...
Searching...
No Matches
Deep Dream

What Deep Dream Is

Deep Dream is an optional ACMXVK neural image-processing stage. It uses a pretrained feature network as a fixed visual vocabulary: ACMXVK selects one activation inside that network, asks LibTorch autograd how the input pixels would need to change to strengthen it, and applies that gradient to the image. The network is not trained and its weights are never modified. Only the current frame is optimized.

This produces structures suggested by the selected model and layer: early features emphasize edges, colors, and fine texture, while deeper features tend to create larger and more semantic forms. The result is still an ordinary ACMX frame and can continue through fragment shaders, compute shaders, multipass chains, texture history, audio reactivity, playlists, crossfades, 3D model texturing, snapshots, and recording.

Deep Dream belongs to the ACMXVK Vulkan backend. It is enabled at build time with -DWITH_DEEP_DREAM=ON and requires CUDA-enabled LibTorch on an NVIDIA GPU. It is separate from OpenCV DNN inference and from the optional acidcam-gpu CUDA filter library.

Video Demonstration

Processing Order

The default frame order is:

  1. Decode or capture the source frame.
  2. Rotate and perform any required source conversion.
  3. Run Deep Dream gradient ascent.
  4. Run the optional acidcam-gpu filter chain.
  5. Upload or hand the result to MXVK.
  6. Run the active Vulkan fragment/compute shader chain.
  7. Update history, texture the optional 3D model, and render output.
  8. Read back for snapshots or recording when required.

With –gpu-filter-before-dream, steps 3 and 4 are reversed. This makes the neural network interpret shapes and colors that acidcam-gpu has already warped. Both orders remain CUDA-resident when compatible CUDA capture, MXVK, OpenCV, and acidcam-gpu builds are present.

Deep Dream does not replace shaders. It preprocesses the texture consumed by the existing Vulkan chain. Shader time continues to follow the configured media clock, so offline rendering speed does not change the animation timing.

Features

The implementation provides:

  • Camera, video-file, and still-image input.
  • VGG16 and Inception V3 feature-model exports.
  • Named or numbered feature-layer selection.
  • All-channel or individual zero-based channel objectives.
  • Configurable gradient iterations and strength.
  • Temporal feedback with independent zoom and rotation.
  • Progressive multi-octave detail restoration.
  • Deterministic spatial jitter based on media-frame sequence.
  • CUDA gradient smoothing for broader structures.
  • A bounded neural working resolution independent of render resolution.
  • Optional FP16 model and working tensors with FP32 loss normalization.
  • Media-clock-based randomized strength, feedback, zoom, rotation, octave count, and octave scale.
  • Traditional independent-frame video rendering in windowed or headless mode.
  • Optional acidcam-gpu processing before or after the neural stage.
  • Direct CUDA capture/NVDEC to LibTorch to Vulkan processing when supported by MXVK, with a compatible host path otherwise.
  • Live settings publication from the Qt interface without restarting ACMXVK.
  • Model, layer, and precision replacement between frames.
  • Safe no-op handling when a finite selected ReLU channel has zero gradient.
  • Frame fallback for runtime neural errors so the capture and recording remain active after Deep Dream is disabled.

HDR frames currently use an RGBA8 compatibility copy for Deep Dream and are restored to the high-precision ACMXVK pipeline afterward. OpenCV DNN effects, still images, HDR input, and maximize-FPS camera mode use the established host path instead of the direct CUDA/Vulkan handoff.

Supported Models and Layers

ACMXVK does not accept an arbitrary classifier directly. It loads an exported TorchScript module with the ACMXVK Deep Dream metadata contract and a forward() method returning an ordered tensor list. The supplied Python exporter creates that format and a readable .pt.json interface sidecar.

Architecture Exported endpoints Default Character
VGG16 relu1_1 through relu5_3 relu4_2 Direct progression from fine textures to larger semantic patterns
Inception V3 Mixed_5b through Mixed_7c Mixed_6c Multi-scale features with varied, complex structures

For Inception V3, Mixed_6c is a strong general starting point. Later blocks such as Mixed_7c produce coarser semantic forms and are more likely to contain inactive individual channels. Use all-channel mode first, then explore specific channels.

Warning
TorchScript files contain executable model code. Load models only from sources you trust.

Building ACMXVK with Deep Dream

Current Platform Support

The primary tested configuration is native Linux with an NVIDIA CUDA GPU. Deep Dream is unavailable on macOS because CUDA is unavailable. The portable Pcons configuration does not currently build this optional feature; use CMake.

Deep Dream and acidcam-gpu use separate switches:

CMake setting Result
WITH_DEEP_DREAM=OFF No LibTorch dependency and no Deep Dream CLI/interface processing
WITH_DEEP_DREAM=ON, WITH_CUDA=OFF Deep Dream enabled; acidcam-gpu filters not linked
WITH_DEEP_DREAM=ON, WITH_CUDA=ON Deep Dream and acidcam-gpu enabled, including selectable processing order

WITH_CUDA in ACMXVK means the acidcam-gpu filter library. Deep Dream still requires CUDA even when that switch is OFF.

Required Development Components

  • CMake 3.20 or newer and a C++20 compiler.
  • Vulkan 1.4 development files and glslc.
  • SDL3, SDL3_ttf, FFmpeg, PNG, Zlib, glm, and the normal ACMXVK dependencies.
  • A working NVIDIA driver and NVIDIA CUDA Toolkit.
  • cuDNN compatible with the selected CUDA/LibTorch build.
  • CUDA-enabled LibTorch containing TorchConfig.cmake.
  • CUDA-enabled OpenCV with core, imgproc, cudaarithm, and cudawarping.
  • MXVK built with CV=ON. Build MXVK with WITH_CUDA=ON for direct CUDA/Vulkan texture interoperability.
  • Torchvision only when exporting VGG16 or Inception V3 models.
  • The installed acidcam-gpu CMake package only when ACMXVK itself is configured with -DWITH_CUDA=ON.

The CUDA Toolkit, cuDNN, LibTorch, OpenCV, and driver versions must be mutually compatible. A CPU-only PyTorch/LibTorch installation cannot be used.

Arch Linux Packages

On an Arch Linux system using CUDA-enabled repository or locally supplied packages, install the equivalent of:

sudo pacman -S --needed base-devel cmake ninja cuda cudnn opencv-cuda \
python-pytorch-cuda python-torchvision-cuda

Package availability and names can vary by configured repositories. The important checks are that OpenCV supplies its CUDA modules and that TorchConfig.cmake describes a CUDA-enabled LibTorch build.

Build and Install MXVK

Build MXVK first. Replace the source and installation paths as appropriate:

cmake -S /path/to/MXVK -B /path/to/MXVK/build-deep-dream \
-G Ninja \
-DCMAKE_BUILD_TYPE=Release \
-DCMAKE_INSTALL_PREFIX=/usr/local \
-DCV=ON \
-DWITH_CUDA=ON
cmake --build /path/to/MXVK/build-deep-dream --parallel
sudo cmake --install /path/to/MXVK/build-deep-dream

WITH_CUDA=AUTO is MXVK's default and enables CUDA when detected, but ON makes a missing CUDA dependency a configuration error rather than silently producing a non-CUDA build. The CMake configure output should state that CUDA and OpenCV support are enabled.

Build ACMXVK

If the CUDA-enabled LibTorch distribution is installed at /opt/libtorch, configure ACMXVK with:

TORCH_CUDA_ARCH_LIST=7.5 cmake -S ACMXVK -B build/acmxvk-dream \
-G Ninja \
-DCMAKE_BUILD_TYPE=Release \
-DCMAKE_PREFIX_PATH=/usr/local \
-DWITH_DEEP_DREAM=ON \
-DWITH_CUDA=OFF \
-DTorch_DIR=/opt/libtorch/share/cmake/Torch
cmake --build build/acmxvk-dream --parallel 2

Use the target GPU's compute capability for TORCH_CUDA_ARCH_LIST. The RTX 2070 uses 7.5. It may be omitted to allow PyTorch to autodetect the available GPU. PyTorch's CMake package may warn that it ignores CMAKE_CUDA_ARCHITECTURES; setting TORCH_CUDA_ARCH_LIST is the supported PyTorch control.

When a system package installs TorchConfig.cmake in a standard CMake location, omit Torch_DIR. For multiple dependency prefixes, quote a semicolon-separated value, for example:

cmake -S ACMXVK -B build/acmxvk-dream \
-DCMAKE_PREFIX_PATH="/usr/local;/opt/libtorch" \
-DWITH_DEEP_DREAM=ON

Add -DWITH_CUDA=ON only to enable acidcam-gpu filters as well. That requires an installed compatible acidcam-gpuConfig.cmake package.

A successful configuration includes a line similar to:

Deep Dream: ENABLED (CUDA LibTorch 2.x)

If the installed executable cannot locate LibTorch at runtime, register /opt/libtorch/lib with the system dynamic loader or provide it through LD_LIBRARY_PATH for the invocation. The dynamic loader solution is preferable for a permanent installation.

Verify the Build

First verify LibTorch, CUDA, cuDNN, and optional MXVK interop without a model:

./build/acmxvk-dream/acmxvk --check-deep-dream --cuda-device 0

Expected output includes:

  • Deep Dream: enabled
  • A LibTorch version.
  • At least one visible CUDA device.
  • CUDA autograd ready on the selected device.
  • cuDNN ready.
  • The CUDA/Vulkan interop status.

Then probe a real exported model. This performs model validation and actual forward/backward gradient-ascent and feedback steps:

./build/acmxvk-dream/acmxvk --check-deep-dream \
--dream-model models/deep-dream-inception-v3.pt \
--dream-layer Mixed_6c \
--dream-channel all \
--cuda-device 0

Run the configured tests with:

ctest --test-dir build/acmxvk-dream --output-on-failure

Exporting Models

The exporter requires Python, PyTorch, and Torchvision. Pretrained weights may be downloaded to PyTorch's user cache on the first run.

Export VGG16:

python ACMXVK/scripts/export_deep_dream_model.py \
--output models/deep-dream-vgg16.pt

Export Inception V3 with Mixed_6c as its default:

python ACMXVK/scripts/export_deep_dream_model.py \
--architecture inception_v3 \
--output models/deep-dream-inception-v3.pt

Use –layers with a comma-separated list to create a smaller model and –default-layer to change its initial endpoint. The default pretrained weights are required for useful visual features. –weights none is for structural/export testing only. Existing output is protected unless –force is supplied.

Keep the generated model.pt.json beside model.pt. ACMXVK reads authoritative metadata embedded in the TorchScript file; the Qt interface uses the sidecar to populate its feature-layer list before launch.

Runtime Options

Option Meaning Valid values and default
–dream-model file.pt Exported TorchScript feature model Required to process frames
–dream-layer name|N Named endpoint or zero-based output index Model default
–dream-iterations N Ascent steps per octave and source frame 1-100; default 1
–dream-strength N Gradient step size Greater than 0 through 10; default 0.05
–dream-feedback N Previous dreamed-frame blend 0-0.99; default 0.9
–dream-zoom N Feedback scale per source frame 0.9-1.1; default 1.01
–dream-rotation N Feedback rotation in degrees per source frame -5 through 5; default 0.1
–dream-size N Maximum neural working dimension 0 or 64-4096; default 512
–dream-fp16 Half-precision model and tensors Disabled by default
–dream-channel N|all One zero-based feature channel or the full activation all, equivalent to -1
–dream-octaves N Progressive image scales 1-8; default 1
–dream-octave-scale N Ratio between adjacent octaves 1.1-3.0; default 1.4
–dream-jitter N Maximum deterministic x/y input shift 0-64 pixels; default 0
–dream-smoothing N Input-gradient box-filter radius 0-16 pixels; default 0
–random-dream seconds Random-setting interval on the media clock Positive seconds
–gpu-filter-before-dream Reverse Deep Dream and acidcam-gpu order Disabled by default
–dream-headless Independent-frame headless video mode Requires input, output, and headless mode
–deep-orig Independent-frame windowed video mode Requires input and output

The legacy spelling –random_dream is accepted. Independent-frame modes cannot be combined with random dream because they intentionally disable feedback animation.

Layers and Channels

All-channel mode maximizes the complete activation and is the most reliable starting point:

--dream-channel all

–dream-channel -1 is equivalent. A nonnegative number selects one channel. The model probe prints the available channel count. Individual channels are exploratory and are not given semantic labels such as "face" or "eye" by ImageNet models.

ReLU channels may be inactive for a particular image. ACMXVK treats a finite zero-gradient iteration as a no-op and tries again on later frames. It does not disable Deep Dream. Non-finite data and genuine CUDA/model errors remain errors.

Feedback Animation

Feedback combines the previous dreamed result with the new source frame. Zoom and rotation transform that previous result before blending it. Reflected borders avoid introducing black edges. Feedback zero makes every frame independent; larger values produce stronger temporal persistence.

Random Dream changes every safe numeric control that was not explicitly specified on the command line: iterations, strength, zoom, rotation, working size, octaves, octave scale, jitter, and smoothing. Explicit values are locked; for example, –dream-jitter 12 keeps jitter at 12 while the remaining unspecified controls change. Temporal feedback, model, layer, channel, precision mode, and processing order remain unchanged. Video uses decoded media time, so an offline render produces the same random-change timing regardless of processing throughput.

Octaves, Jitter, and Smoothing

Multiple octaves begin at a smaller image and advance toward the configured working size. Source detail is restored between scales. Cost grows roughly as the number of unique octave sizes multiplied by the iteration count.

Jitter rolls the input before each model pass and maps gradients back through that operation. It reduces stationary grid and edge bias without adding a second network pass. Gradient smoothing uses a CUDA box filter to encourage wider coherent features. Radius 1 or 2 is a useful starting point.

Resolution and Precision

The default –dream-size 512 bounds neural work while the Vulkan pipeline and recording keep the original output resolution. Use a smaller value for faster previews and a larger value for finer structure. The working width and height must both remain at least the model minimum: 32 pixels for the full VGG16 export and 75 pixels for the full Inception V3 export. Portrait or widescreen aspect ratios can therefore require a longest dimension greater than that minimum.

FP16 reduces model/tensor memory and can improve speed on supported NVIDIA hardware. Loss and gradient magnitude are still evaluated in FP32. Use FP32 if the selected GPU has poor half-precision performance or a model exhibits numerical instability.

Examples

Real-Time Video with Vulkan Shaders

./build/acmxvk-dream/acmxvk \
--input input.mp4 \
--dream-model models/deep-dream-inception-v3.pt \
--dream-layer Mixed_6c \
--dream-channel all \
--dream-iterations 1 \
--dream-strength 0.05 \
--dream-size 512 \
--dream-fp16 \
--shaders shaders_acmxvk \
--shader-file color-effect.frag.spv

Camera with Feedback

./build/acmxvk-dream/acmxvk \
--device 0 \
--dream-model models/deep-dream-inception-v3.pt \
--dream-layer Mixed_6c \
--dream-feedback 0.85 \
--dream-zoom 1.01 \
--dream-rotation 0.15 \
--dream-size 512 \
--shaders shaders_acmxvk

Randomized Headless Visuals

./build/acmxvk-dream/acmxvk \
--input input.mp4 --output randomized.mp4 --headless \
--dream-model models/deep-dream-vgg16.pt \
--dream-layer relu4_2 \
--random-dream 0.5 \
--shaders shaders_acmxvk \
--shader-file color-effect.frag.spv

Traditional Independent-Frame Video

Use –dream-headless for offline recording without a window:

./build/acmxvk-dream/acmxvk \
--input input.mp4 --output dreamed.mp4 \
--headless --dream-headless \
--dream-model models/deep-dream-vgg16.pt \
--dream-layer relu4_2 \
--dream-iterations 4 \
--dream-octaves 3 \
--dream-strength 0.04

This forces feedback to 0, zoom to 1, rotation to 0, and enables no-drop encoding. Every decoded frame is processed independently. Use –deep-orig instead of –headless –dream-headless for the same processing with a preview window.

acidcam-gpu Before Deep Dream

This configuration requires both build switches and an acidcam-gpu filter:

./build/acmxvk-dream/acmxvk \
--input input.mp4 \
--gpu-filter 252,608,330 --gpu-buffer 8 \
--gpu-filter-before-dream \
--dream-model models/deep-dream-inception-v3.pt \
--dream-layer Mixed_6c \
--dream-channel all \
--shaders shaders_acmxvk

Qt Interface

When the interface detects a Deep Dream-enabled ACMXVK executable, open Session > Deep Dream Settings.

  • Enable Deep Dream and browse to the exported .pt model.
  • Keep its .pt.json sidecar beside it so the layer selector is populated with the exact model endpoints.
  • Select a layer and configure ascent, feedback, detail, and performance.
  • Choose All channels for the complete activation.
  • Enable acidcam-gpu before Deep Dream when a CUDA filter chain should preprocess the neural input.
  • Use Apply while ACMXVK is running to publish settings immediately.
  • Use Randomize to create and apply bounded random settings.

The dialog is modeless, so the GPU filter and other settings dialogs can remain open. Scalar changes apply between frames. A model, layer, or precision change constructs and validates a replacement model before swapping it into the render path. Invalid live settings leave the current working configuration active.

The Original/independent-frame interface option adds –deep-orig. It requires video input and recording output. The window remains visible, but feedback zoom and rotation are disabled for traditional frame-by-frame Deep Dream processing.

Performance Guidance

Start with:

iterations: 1
strength: 0.03 to 0.05
channel: all
octaves: 1
dream size: 384 or 512
jitter: 0 to 8
smoothing: 0 or 1
FP16: enabled on suitable NVIDIA GPUs

The largest costs are working resolution, iteration count, and octave count. Doubling image dimensions approaches four times the neural pixel work. Five iterations over six octaves can be dramatically slower than the one-iteration, one-octave real-time default. Output resolution does not need to equal neural resolution: ACMXVK restores the dreamed result to the source-sized Vulkan texture before shaders and encoding.

Deeper layers create different imagery rather than automatically improving quality. If a specific channel appears inactive, return to all-channel mode or try another channel/layer. FP32 is useful when diagnosing FP16 behavior.

Runtime Safety and Recovery

ACMXVK validates the model format, version, normalization, layer table, output count, tensor shapes, selected layer, and channel range before media processing. Each frame validates activation loss, gradients, and output pixels.

A finite zero gradient is valid for a dead ReLU channel and causes a no-op iteration. If a genuine per-frame neural exception occurs, ACMXVK reports it, disables Deep Dream, and sends the current source or acidcam-gpu frame through the remaining Vulkan pipeline. Capture and recording continue. This fallback prevents a neural failure from being mistaken for video end-of-file.

Troubleshooting

CMake cannot find TorchConfig.cmake

Set Torch_DIR to the directory containing that file, commonly /opt/libtorch/share/cmake/Torch, or add the LibTorch prefix to CMAKE_PREFIX_PATH.

WITH_DEEP_DREAM requires a CUDA-enabled LibTorch installation

The selected Torch package is CPU-only or its CUDA dependencies were not found. Install a CUDA build matching the system toolkit and driver, then delete the CMake cache or configure a fresh build directory.

LibTorch CUDA devices: 0

Confirm nvidia-smi works. In a container, recreate or start it with NVIDIA device access; installing CUDA libraries inside a container does not automatically expose the host GPU.

PyTorch ignores CMAKE_CUDA_ARCHITECTURES

This warning comes from PyTorch's CMake configuration. Use the environment variable TORCH_CUDA_ARCH_LIST, for example 7.5 for an RTX 2070.

library kineto not found

This is normally a nonfatal warning from a LibTorch package built without the optional Kineto profiler. ACMXVK Deep Dream does not require Kineto.

Deep Dream working image is smaller than the model minimum

Increase –dream-size. Both dimensions after aspect-ratio-preserving resize must be at least 32 for VGG16 or 75 for the full Inception V3 export.

An individual Inception channel sometimes has no effect

That channel may be inactive after its ReLU for the current frame. Use –dream-channel all, another channel, or another layer. ACMXVK safely skips zero-gradient iterations and keeps the feature enabled.

The program cannot load a LibTorch shared library

Add the LibTorch lib directory to the system dynamic-loader configuration or supply it through LD_LIBRARY_PATH. Re-run –check-deep-dream after correcting the loader path.

Out of memory or very slow processing

Lower –dream-size, iterations, or octaves; enable FP16 on appropriate hardware; and use one feature channel only after confirming it is active. Close other GPU-intensive applications when VRAM is constrained.

See ACMXVK Vulkan Backend for the complete Vulkan backend architecture and ACMXVK/README.md for the evolving implementation reference.