Inputs
Note: The latent tensor is spatially downsampled by a factor of 8 in both height and width, and contains 16 channels. The number of latent temporal frames is calculated as
((length - 1) // 8) + 1.
Outputs
This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub
Source fingerprint (SHA-256):
7ee194324b02367ed853f6d36bc51742081bac6a9469c4a619586e0560a1b33b