Deep Dream is an optional ACMXVK neural image-processing stage. It uses a pretrained feature network as a fixed visual vocabulary: ACMXVK selects one activation inside that network, asks LibTorch autograd how the input pixels would need to change to strengthen it, and applies that gradient to the image. The network is not trained and its weights are never modified. Only the current frame is optimized.
This produces structures suggested by the selected model and layer: early features emphasize edges, colors, and fine texture, while deeper features tend to create larger and more semantic forms. The result is still an ordinary ACMX frame and can continue through fragment shaders, compute shaders, multipass chains, texture history, audio reactivity, playlists, crossfades, 3D model texturing, snapshots, and recording.
Deep Dream belongs to the ACMXVK Vulkan backend. It is enabled at build time with -DWITH_DEEP_DREAM=ON and requires CUDA-enabled LibTorch on an NVIDIA GPU. It is separate from OpenCV DNN inference and from the optional acidcam-gpu CUDA filter library.
The default frame order is:
With –gpu-filter-before-dream, steps 3 and 4 are reversed. This makes the neural network interpret shapes and colors that acidcam-gpu has already warped. Both orders remain CUDA-resident when compatible CUDA capture, MXVK, OpenCV, and acidcam-gpu builds are present.
Deep Dream does not replace shaders. It preprocesses the texture consumed by the existing Vulkan chain. Shader time continues to follow the configured media clock, so offline rendering speed does not change the animation timing.
The implementation provides:
HDR frames currently use an RGBA8 compatibility copy for Deep Dream and are restored to the high-precision ACMXVK pipeline afterward. OpenCV DNN effects, still images, HDR input, and maximize-FPS camera mode use the established host path instead of the direct CUDA/Vulkan handoff.
ACMXVK does not accept an arbitrary classifier directly. It loads an exported TorchScript module with the ACMXVK Deep Dream metadata contract and a forward() method returning an ordered tensor list. The supplied Python exporter creates that format and a readable .pt.json interface sidecar.
| Architecture | Exported endpoints | Default | Character |
|---|---|---|---|
| VGG16 | relu1_1 through relu5_3 | relu4_2 | Direct progression from fine textures to larger semantic patterns |
| Inception V3 | Mixed_5b through Mixed_7c | Mixed_6c | Multi-scale features with varied, complex structures |
For Inception V3, Mixed_6c is a strong general starting point. Later blocks such as Mixed_7c produce coarser semantic forms and are more likely to contain inactive individual channels. Use all-channel mode first, then explore specific channels.
The primary tested configuration is native Linux with an NVIDIA CUDA GPU. Deep Dream is unavailable on macOS because CUDA is unavailable. The portable Pcons configuration does not currently build this optional feature; use CMake.
Deep Dream and acidcam-gpu use separate switches:
| CMake setting | Result |
|---|---|
| WITH_DEEP_DREAM=OFF | No LibTorch dependency and no Deep Dream CLI/interface processing |
| WITH_DEEP_DREAM=ON, WITH_CUDA=OFF | Deep Dream enabled; acidcam-gpu filters not linked |
| WITH_DEEP_DREAM=ON, WITH_CUDA=ON | Deep Dream and acidcam-gpu enabled, including selectable processing order |
WITH_CUDA in ACMXVK means the acidcam-gpu filter library. Deep Dream still requires CUDA even when that switch is OFF.
The CUDA Toolkit, cuDNN, LibTorch, OpenCV, and driver versions must be mutually compatible. A CPU-only PyTorch/LibTorch installation cannot be used.
On an Arch Linux system using CUDA-enabled repository or locally supplied packages, install the equivalent of:
Package availability and names can vary by configured repositories. The important checks are that OpenCV supplies its CUDA modules and that TorchConfig.cmake describes a CUDA-enabled LibTorch build.
Build MXVK first. Replace the source and installation paths as appropriate:
WITH_CUDA=AUTO is MXVK's default and enables CUDA when detected, but ON makes a missing CUDA dependency a configuration error rather than silently producing a non-CUDA build. The CMake configure output should state that CUDA and OpenCV support are enabled.
If the CUDA-enabled LibTorch distribution is installed at /opt/libtorch, configure ACMXVK with:
Use the target GPU's compute capability for TORCH_CUDA_ARCH_LIST. The RTX 2070 uses 7.5. It may be omitted to allow PyTorch to autodetect the available GPU. PyTorch's CMake package may warn that it ignores CMAKE_CUDA_ARCHITECTURES; setting TORCH_CUDA_ARCH_LIST is the supported PyTorch control.
When a system package installs TorchConfig.cmake in a standard CMake location, omit Torch_DIR. For multiple dependency prefixes, quote a semicolon-separated value, for example:
Add -DWITH_CUDA=ON only to enable acidcam-gpu filters as well. That requires an installed compatible acidcam-gpuConfig.cmake package.
A successful configuration includes a line similar to:
If the installed executable cannot locate LibTorch at runtime, register /opt/libtorch/lib with the system dynamic loader or provide it through LD_LIBRARY_PATH for the invocation. The dynamic loader solution is preferable for a permanent installation.
First verify LibTorch, CUDA, cuDNN, and optional MXVK interop without a model:
Expected output includes:
Then probe a real exported model. This performs model validation and actual forward/backward gradient-ascent and feedback steps:
Run the configured tests with:
The exporter requires Python, PyTorch, and Torchvision. Pretrained weights may be downloaded to PyTorch's user cache on the first run.
Export VGG16:
Export Inception V3 with Mixed_6c as its default:
Use –layers with a comma-separated list to create a smaller model and –default-layer to change its initial endpoint. The default pretrained weights are required for useful visual features. –weights none is for structural/export testing only. Existing output is protected unless –force is supplied.
Keep the generated model.pt.json beside model.pt. ACMXVK reads authoritative metadata embedded in the TorchScript file; the Qt interface uses the sidecar to populate its feature-layer list before launch.
| Option | Meaning | Valid values and default |
|---|---|---|
| –dream-model file.pt | Exported TorchScript feature model | Required to process frames |
| –dream-layer name|N | Named endpoint or zero-based output index | Model default |
| –dream-iterations N | Ascent steps per octave and source frame | 1-100; default 1 |
| –dream-strength N | Gradient step size | Greater than 0 through 10; default 0.05 |
| –dream-feedback N | Previous dreamed-frame blend | 0-0.99; default 0.9 |
| –dream-zoom N | Feedback scale per source frame | 0.9-1.1; default 1.01 |
| –dream-rotation N | Feedback rotation in degrees per source frame | -5 through 5; default 0.1 |
| –dream-size N | Maximum neural working dimension | 0 or 64-4096; default 512 |
| –dream-fp16 | Half-precision model and tensors | Disabled by default |
| –dream-channel N|all | One zero-based feature channel or the full activation | all, equivalent to -1 |
| –dream-octaves N | Progressive image scales | 1-8; default 1 |
| –dream-octave-scale N | Ratio between adjacent octaves | 1.1-3.0; default 1.4 |
| –dream-jitter N | Maximum deterministic x/y input shift | 0-64 pixels; default 0 |
| –dream-smoothing N | Input-gradient box-filter radius | 0-16 pixels; default 0 |
| –random-dream seconds | Random-setting interval on the media clock | Positive seconds |
| –gpu-filter-before-dream | Reverse Deep Dream and acidcam-gpu order | Disabled by default |
| –dream-headless | Independent-frame headless video mode | Requires input, output, and headless mode |
| –deep-orig | Independent-frame windowed video mode | Requires input and output |
The legacy spelling –random_dream is accepted. Independent-frame modes cannot be combined with random dream because they intentionally disable feedback animation.
All-channel mode maximizes the complete activation and is the most reliable starting point:
–dream-channel -1 is equivalent. A nonnegative number selects one channel. The model probe prints the available channel count. Individual channels are exploratory and are not given semantic labels such as "face" or "eye" by ImageNet models.
ReLU channels may be inactive for a particular image. ACMXVK treats a finite zero-gradient iteration as a no-op and tries again on later frames. It does not disable Deep Dream. Non-finite data and genuine CUDA/model errors remain errors.
Feedback combines the previous dreamed result with the new source frame. Zoom and rotation transform that previous result before blending it. Reflected borders avoid introducing black edges. Feedback zero makes every frame independent; larger values produce stronger temporal persistence.
Random Dream changes every safe numeric control that was not explicitly specified on the command line: iterations, strength, zoom, rotation, working size, octaves, octave scale, jitter, and smoothing. Explicit values are locked; for example, –dream-jitter 12 keeps jitter at 12 while the remaining unspecified controls change. Temporal feedback, model, layer, channel, precision mode, and processing order remain unchanged. Video uses decoded media time, so an offline render produces the same random-change timing regardless of processing throughput.
Multiple octaves begin at a smaller image and advance toward the configured working size. Source detail is restored between scales. Cost grows roughly as the number of unique octave sizes multiplied by the iteration count.
Jitter rolls the input before each model pass and maps gradients back through that operation. It reduces stationary grid and edge bias without adding a second network pass. Gradient smoothing uses a CUDA box filter to encourage wider coherent features. Radius 1 or 2 is a useful starting point.
The default –dream-size 512 bounds neural work while the Vulkan pipeline and recording keep the original output resolution. Use a smaller value for faster previews and a larger value for finer structure. The working width and height must both remain at least the model minimum: 32 pixels for the full VGG16 export and 75 pixels for the full Inception V3 export. Portrait or widescreen aspect ratios can therefore require a longest dimension greater than that minimum.
FP16 reduces model/tensor memory and can improve speed on supported NVIDIA hardware. Loss and gradient magnitude are still evaluated in FP32. Use FP32 if the selected GPU has poor half-precision performance or a model exhibits numerical instability.
Use –dream-headless for offline recording without a window:
This forces feedback to 0, zoom to 1, rotation to 0, and enables no-drop encoding. Every decoded frame is processed independently. Use –deep-orig instead of –headless –dream-headless for the same processing with a preview window.
This configuration requires both build switches and an acidcam-gpu filter:
When the interface detects a Deep Dream-enabled ACMXVK executable, open Session > Deep Dream Settings.
The dialog is modeless, so the GPU filter and other settings dialogs can remain open. Scalar changes apply between frames. A model, layer, or precision change constructs and validates a replacement model before swapping it into the render path. Invalid live settings leave the current working configuration active.
The Original/independent-frame interface option adds –deep-orig. It requires video input and recording output. The window remains visible, but feedback zoom and rotation are disabled for traditional frame-by-frame Deep Dream processing.
Start with:
The largest costs are working resolution, iteration count, and octave count. Doubling image dimensions approaches four times the neural pixel work. Five iterations over six octaves can be dramatically slower than the one-iteration, one-octave real-time default. Output resolution does not need to equal neural resolution: ACMXVK restores the dreamed result to the source-sized Vulkan texture before shaders and encoding.
Deeper layers create different imagery rather than automatically improving quality. If a specific channel appears inactive, return to all-channel mode or try another channel/layer. FP32 is useful when diagnosing FP16 behavior.
ACMXVK validates the model format, version, normalization, layer table, output count, tensor shapes, selected layer, and channel range before media processing. Each frame validates activation loss, gradients, and output pixels.
A finite zero gradient is valid for a dead ReLU channel and causes a no-op iteration. If a genuine per-frame neural exception occurs, ACMXVK reports it, disables Deep Dream, and sends the current source or acidcam-gpu frame through the remaining Vulkan pipeline. Capture and recording continue. This fallback prevents a neural failure from being mistaken for video end-of-file.
CMake cannot find TorchConfig.cmake
Set Torch_DIR to the directory containing that file, commonly /opt/libtorch/share/cmake/Torch, or add the LibTorch prefix to CMAKE_PREFIX_PATH.
WITH_DEEP_DREAM requires a CUDA-enabled LibTorch installation
The selected Torch package is CPU-only or its CUDA dependencies were not found. Install a CUDA build matching the system toolkit and driver, then delete the CMake cache or configure a fresh build directory.
LibTorch CUDA devices: 0
Confirm nvidia-smi works. In a container, recreate or start it with NVIDIA device access; installing CUDA libraries inside a container does not automatically expose the host GPU.
PyTorch ignores CMAKE_CUDA_ARCHITECTURES
This warning comes from PyTorch's CMake configuration. Use the environment variable TORCH_CUDA_ARCH_LIST, for example 7.5 for an RTX 2070.
library kineto not found
This is normally a nonfatal warning from a LibTorch package built without the optional Kineto profiler. ACMXVK Deep Dream does not require Kineto.
Deep Dream working image is smaller than the model minimum
Increase –dream-size. Both dimensions after aspect-ratio-preserving resize must be at least 32 for VGG16 or 75 for the full Inception V3 export.
An individual Inception channel sometimes has no effect
That channel may be inactive after its ReLU for the current frame. Use –dream-channel all, another channel, or another layer. ACMXVK safely skips zero-gradient iterations and keeps the feature enabled.
The program cannot load a LibTorch shared library
Add the LibTorch lib directory to the system dynamic-loader configuration or supply it through LD_LIBRARY_PATH. Re-run –check-deep-dream after correcting the loader path.
Out of memory or very slow processing
Lower –dream-size, iterations, or octaves; enable FP16 on appropriate hardware; and use one feature channel only after confirming it is active. Close other GPU-intensive applications when VRAM is constrained.
See ACMXVK Vulkan Backend for the complete Vulkan backend architecture and ACMXVK/README.md for the evolving implementation reference.