|
Download API.md from data-archetype/canter: direct link, hf CLI and curl.
- Browser
- Download file 24.2 kB
-
https://huggingface.co/data-archetype/canter/resolve/main/API.md
- Command line
-
hf download hf://data-archetype/canter/API.md
-
curl -L -o API.md https://huggingface.co/data-archetype/canter/resolve/main/API.md
24.2 kB
| # Canter API and inference parameters | |
| [Model card](README.md) · [Example gallery](GALLERY.md) · | |
| [Technical report](TECHNICAL_REPORT.md) | |
| Canter exposes a high-level image pipeline and a lower-level latent inference | |
| engine. Configuration uses frozen dataclasses and enums. | |
| The installed `canter` package always supplies the inference implementation. | |
| For remote model IDs, `revision` selects checkpoint artifacts only; selecting | |
| an older checkpoint does not load or execute the Python package bundled in | |
| that historical repository snapshot. | |
| ## Minimal use | |
| ```python | |
| from canter import CanterPipeline | |
| pipe = CanterPipeline.from_pretrained("data-archetype/canter") | |
| output = pipe( | |
| "A red rally car driving around a wet forest road", | |
| negative_prompts="Compression artefacts", | |
| ) | |
| image = output.image | |
| ``` | |
| ### Generation arguments | |
| | Parameter | Default | Description | | |
| | --- | --- | --- | | |
| | `prompts` | required | One prompt string or a sequence containing one prompt per image. | | |
| | `negative_prompts` | `None` | Optional negative prompt string or sequence for CFG and contrastive PDG. The count must match `prompts`. `None` uses the learned unconditional token. | | |
| | `config` | `CanterPipelineConfig()` | Inference and output settings. | | |
| | `initial_noise` | `None` | Optional float32 latent noise tensor with the configured batch and spatial shape. | | |
| | `progress` | `None` | Optional callback receiving completed and total solver updates. | | |
| | `preview` | `None` | Optional callback receiving asynchronous latent-RGB PIL images plus completed and total solver updates. | | |
| `CanterPipeline` loads the flow-matching denoiser and text tokenizer with the | |
| bundled | |
| [`SmolLM2-360M`](https://huggingface.co/HuggingFaceTB/SmolLM2-360M) weights. It | |
| compiles the selected text backend, downloads DINAC-AE-D2, generates latents, | |
| and decodes them into images. | |
| ## Loading the pipeline | |
| ```python | |
| from canter import CanterPipeline, TextAttentionBackend, WeightDType | |
| pipe = CanterPipeline.from_pretrained( | |
| "data-archetype/canter", | |
| dtype=WeightDType.BFLOAT16, | |
| text_backend=TextAttentionBackend.DENSE, | |
| device="cuda", | |
| revision=None, | |
| cache_dir=None, | |
| compile_model=True, | |
| ) | |
| ``` | |
| ### `CanterPipeline.from_pretrained` | |
| | Parameter | Default | Description | | |
| | --- | --- | --- | | |
| | `path_or_repo_id` | required | A Hugging Face repository ID or downloaded model repository directory. | | |
| | `dtype` | `WeightDType.BFLOAT16` | Weight storage dtype. Compute still uses bfloat16 AMP with explicit float32 operations. | | |
| | `text_backend` | `TextAttentionBackend.DENSE` | Bundled text-encoder, text-refinement, and cross-attention layout. | | |
| | `device` | `"cuda"` | CUDA device string or `torch.device`. | | |
| | `revision` | `None` | Hugging Face checkpoint branch, commit, or immutable release tag. `None` uses the installed package's pinned default checkpoint. An explicit tag such as `"v0001"` loads those weights with the currently installed code. | | |
| | `cache_dir` | `None` | Optional Hugging Face cache directory. | | |
| | `compile_model` | `True` | Compile the selected inference kernels. Compilation errors are reported directly. | | |
| | `vae` | `None` | Optional compatible `CanterVae` instance. `None` downloads DINAC-AE-D2 automatically. | | |
| ### Weight dtypes | |
| | Enum | Value | Description | | |
| | --- | --- | --- | | |
| | `WeightDType.BFLOAT16` | `"bfloat16"` | Default compact release with required float32 parameter islands retained. | | |
| | `WeightDType.FLOAT32` | `"float32"` | Full-float32 weight storage release. Model compute remains bfloat16 AMP. | | |
| ### Text backends | |
| | Enum | Value | Description | | |
| | --- | --- | --- | | |
| | `TextAttentionBackend.DENSE` | `"dense"` | Compatibility default using padded tensors and standard PyTorch scaled dot-product attention throughout text encoding and denoising. | | |
| | `TextAttentionBackend.JAGGED` | `"jagged"` | Explicit optimization using packed text, CUDA FlashAttention, and jagged NestedTensor kernels throughout text encoding and denoising. | | |
| Only the selected backend is compiled. The backend cannot be changed after | |
| compilation. | |
| ## Pipeline configuration | |
| ```python | |
| from canter import ( | |
| CanterInferenceConfig, | |
| CanterOutputType, | |
| CanterPipelineConfig, | |
| ) | |
| config = CanterPipelineConfig( | |
| inference=CanterInferenceConfig(), | |
| output_type=CanterOutputType.PIL, | |
| ) | |
| output = pipe("A portrait lit by a large north-facing window", config=config) | |
| ``` | |
| ### `CanterPipelineConfig` | |
| | Field | Default | Description | | |
| | --- | --- | --- | | |
| | `inference` | `CanterInferenceConfig()` | Latent generation settings. | | |
| | `output_type` | `CanterOutputType.PIL` | Selected decoded or latent output representation. | | |
| ### Output types | |
| | Enum | Result | | |
| | --- | --- | | |
| | `CanterOutputType.PIL` | `output.images` contains one PIL image per prompt. `output.image` returns the sole image for batch size one. | | |
| | `CanterOutputType.TENSOR` | `output.image_tensor` contains clamped float32 RGB pixels in `[-1, 1]`. | | |
| | `CanterOutputType.LATENT` | VAE decoding is skipped. `output.latents` contains whitened float32 model latents. | | |
| Every output also contains the descending float32 solver schedule in | |
| `output.schedule`. | |
| ## Live latent previews | |
| ```python | |
| def show_preview(images, completed, total): | |
| images[0].show() | |
| output = pipe( | |
| "A lighthouse in a winter storm", | |
| preview=show_preview, | |
| ) | |
| ``` | |
| The bundled float32 linear projection maps the whitened 128-channel DINAC | |
| latent directly to RGB at one eighth of the requested image size. Every solver | |
| update is eligible by default. If the previous preview is still transferring, | |
| converting, or running the callback, the new update is dropped instead of | |
| blocking sampling. The callback receives this native one-eighth-size image; | |
| the web UI lets the browser scale it for display. | |
| Accepted previews perform only a 1×1 projection and pixel shuffle on CUDA. | |
| The small RGB result is copied non-blockingly into pinned host memory. CUDA | |
| event waiting, PIL conversion, and the application callback all run on a | |
| single background worker. Every step is eligible; the one-worker busy-drop | |
| policy supplies backpressure without a separate callback cap. | |
| ## Inference configuration | |
| ### Recommended defaults | |
| The `CanterInferenceConfig` defaults are also the recommended starting point | |
| for this release. They use 50 ER-SDE updates with noise scale `1.0`, a linear | |
| schedule with log-SNR shift `-2.3`, Contrastive PDG at scale `2.5`, disabled | |
| CFG, and main-path self-attention gain `-0.05`. | |
| ```python | |
| from canter import ( | |
| CanterInferenceConfig, | |
| CfgGuidance, | |
| PdgCurve, | |
| PdgGuidance, | |
| PdgMode, | |
| Schedule, | |
| Solver, | |
| ) | |
| config = CanterInferenceConfig( | |
| height=1216, | |
| width=832, | |
| steps=50, | |
| solver=Solver.ER_SDE, | |
| schedule=Schedule.LINEAR, | |
| log_snr_shift=-2.3, | |
| cfg=CfgGuidance( | |
| enabled=False, | |
| scale=None, | |
| start_step=0, | |
| stop_step=None, | |
| ), | |
| pdg=PdgGuidance( | |
| enabled=True, | |
| mode=PdgMode.FULL_CONTRASTIVE, | |
| curve=PdgCurve.CONSTANT, | |
| noisy_scale=2.5, | |
| clean_scale=2.5, | |
| power=3.0, | |
| start_step=0, | |
| stop_step=None, | |
| ), | |
| self_attention_gain=-0.05, | |
| sde_noise_multiplier=1.0, | |
| seed=42, | |
| generator=None, | |
| ) | |
| ``` | |
| The recommended Contrastive PDG recipe additionally supplies | |
| `Compression artefacts` through the separate `negative_prompts` generation | |
| argument. With Contrastive PDG, this text conditions the middle-skipped weak | |
| path while the positive prompt conditions the full main path. In the tested | |
| release settings it acts as an additional inference regulariser and increases | |
| image quality. Leave it unchanged for the recommended baseline; blank or | |
| omitted negative text instead uses Canter's learned unconditional conditioning | |
| and therefore produces a different guidance trajectory. | |
| Choosing between this Contrastive PDG recipe and standard full-path PDG is | |
| somewhat trial and error. Larger images generally work better with Contrastive | |
| PDG and the default `Compression artefacts` negative text, while standard PDG | |
| can be better for particular prompts or smaller images. ABM2 and ER-SDE tend to | |
| be the strongest sampler choices. At the same update count, ABM2 is | |
| approximately twice as slow because its corrected states require additional | |
| denoiser evaluations. | |
| ### `CanterInferenceConfig` | |
| | Field | Default | Description | | |
| | --- | --- | --- | | |
| | `height` | `1216` | Output height in pixels. Must be positive and divisible by 16. | | |
| | `width` | `832` | Output width in pixels. Must be positive and divisible by 16. | | |
| | `steps` | `50` | Number of solver state updates. A run with 50 updates uses 51 schedule points. | | |
| | `solver` | `Solver.ER_SDE` | Numerical solver. | | |
| | `schedule` | `Schedule.LINEAR` | Timestep spacing. | | |
| | `log_snr_shift` | `-2.3` | Additive log-SNR schedule shift. | | |
| | `cfg` | disabled | Classifier-free guidance settings. | | |
| | `pdg` | Contrastive PDG 2.5 | Path-drop guidance settings. | | |
| | `self_attention_gain` | `-0.05` | Gain applied exclusively to image self-attention on the main denoiser path. | | |
| | `sde_noise_multiplier` | `1.0` | Shared non-negative stochasticity control: DPM++ eta for DPM++ 2M SDE, and noise scale for Euler-Maruyama, ER-SDE, and ER-SDE Brownian tree. Deterministic solvers ignore it. The web UI selects the solver-specific presets described below. | | |
| | `seed` | `42` | Random seed. Set to `None` when supplying `generator`. | | |
| | `generator` | `None` | Optional CUDA `torch.Generator` on the same device as the model. Exactly one of `seed` and `generator` is required. | | |
| The seed controls latent noise, SPRINT routing, and stochastic solver | |
| operations. The VAE decoder retains its published seed behavior. | |
| ### PNG metadata | |
| Images downloaded from the Gradio interface contain a `canter` PNG text field | |
| with compact JSON. The object begins with the prompt, optional negative prompt, | |
| effective per-image seed, width, height, steps, solver, and schedule. It then | |
| records PDG, CFG, | |
| self-attention gain, logSNR shift, the DPM++ eta or SDE noise scale, installed | |
| code version, numbered checkpoint release, and weight dtype. | |
| ### Solvers | |
| | Enum | Value | Description | | |
| | --- | --- | --- | | |
| | `Solver.EULER` | `"euler"` | First-order deterministic Euler updates. | | |
| | `Solver.EULER_MARUYAMA` | `"euler_maruyama"` | Stochastic reverse-SDE updates. Uses `sde_noise_multiplier`. | | |
| | `Solver.ER_SDE` | `"er_sde"` | Third-stage VP ER-SDE with 16-point Gauss-Legendre correction quadrature and independent per-step normals. Uses `sde_noise_multiplier`. | | |
| | `Solver.ER_SDE_BT` | `"er_sde_bt"` | The same ER-SDE update with fixed-domain Brownian-tree increments for nested-grid path coupling. Uses `sde_noise_multiplier`. | | |
| | `Solver.DPMPP_2M` | `"dpmpp_2m"` | Flow-matching DPM++ 2M updates with finite pre-shift endpoint handling. | | |
| | `Solver.DPMPP_2M_SDE` | `"dpmpp_2m_sde"` | Stochastic midpoint DPM++ 2M updates. Uses `sde_noise_multiplier` as eta; zero exactly recovers deterministic DPM++ 2M. | | |
| | `Solver.ABM2` | `"abm2"` | Variable-step Adams-Bashforth-Moulton updates with corrected-state reevaluation. | | |
| ER-SDE is the default. ABM2 performs additional denoiser evaluations for its | |
| corrected states when selected. | |
| Euler-Maruyama, ER-SDE, and DPM++ 2M SDE draw independent normals at each | |
| stochastic update. `Solver.ER_SDE_BT` is the opt-in coupled variant. It uses one | |
| fixed-entropy tree per image on the schedule-independent mathematical ER clock | |
| domain, so subdivided increments add back to their coarse increment across | |
| nested grids. Its integer seeds are reproducible but do not identify the same | |
| trajectory as `Solver.ER_SDE` seeds. | |
| The bundled web interface resets the shared stochasticity control when the | |
| solver changes: DPM++ 2M SDE selects eta `0.5`; every other solver selects | |
| `1.0`, including the ER-SDE, ER-SDE Brownian-tree, and Euler-Maruyama | |
| noise-scale controls. Direct API configurations use the explicit | |
| `sde_noise_multiplier` supplied by the caller. | |
| ### Schedules | |
| | Enum | Value | Description | | |
| | --- | --- | --- | | |
| | `Schedule.LINEAR` | `"linear"` | Uniform spacing from noisy to clean. | | |
| | `Schedule.BETA` | `"beta"` | Beta(0.6, 0.6) quantile spacing with more schedule density near the endpoints. | | |
| Schedules and solver state remain float32. | |
| In the bundled web interface, selecting Beta(0.6, 0.6) resets the log-SNR shift | |
| to `0.0`; selecting Linear resets it to `-2.3`. The shift remains independently | |
| editable after either preset is applied. Direct API configurations use the | |
| explicit `log_snr_shift` supplied by the caller. | |
| ## CFG | |
| ```python | |
| from canter import ( | |
| CanterInferenceConfig, | |
| CanterPipelineConfig, | |
| CfgGuidance, | |
| PdgGuidance, | |
| PdgMode, | |
| ) | |
| cfg = CfgGuidance( | |
| enabled=True, | |
| scale=3.0, | |
| start_step=0, | |
| stop_step=29, | |
| ) | |
| output = pipe( | |
| "A studio portrait with soft natural light", | |
| negative_prompts="oversaturated, harsh contrast", | |
| config=CanterPipelineConfig( | |
| inference=CanterInferenceConfig( | |
| cfg=cfg, | |
| pdg=PdgGuidance(mode=PdgMode.COMBINED_CFG_PDG), | |
| ), | |
| ), | |
| ) | |
| ``` | |
| ### `CfgGuidance` | |
| | Field | Default | Description | | |
| | --- | --- | --- | | |
| | `enabled` | `False` | Enable classifier-free guidance. | | |
| | `scale` | `None` | Non-negative guidance scale. Required when CFG is enabled or when the selected PDG mode uses CFG. | | |
| | `start_step` | `0` | First active solver update, inclusive. | | |
| | `stop_step` | `None` | Last active solver update, inclusive. `None` selects the final update. | | |
| Step indices run from `0` through `steps - 1`. | |
| An explicit negative prompt replaces the learned unconditional token on CFG | |
| and contrastive PDG branches. Negative prompts require enabled CFG or a PDG | |
| mode that consumes contrastive text. For prompt batches, supply one negative | |
| prompt per positive prompt. | |
| ## PDG | |
| ```python | |
| from canter import PdgCurve, PdgGuidance, PdgMode | |
| pdg = PdgGuidance( | |
| enabled=True, | |
| mode=PdgMode.FULL_CONTRASTIVE, | |
| curve=PdgCurve.POWER, | |
| noisy_scale=2.0, | |
| clean_scale=2.5, | |
| power=3.0, | |
| start_step=0, | |
| stop_step=None, | |
| ) | |
| ``` | |
| ### `PdgGuidance` | |
| | Field | Default | Description | | |
| | --- | --- | --- | | |
| | `enabled` | `True` | Enable path-drop guidance. | | |
| | `mode` | `PdgMode.FULL_CONTRASTIVE` | Path and CFG interaction policy. | | |
| | `curve` | `PdgCurve.CONSTANT` | Scale interpolation from noisy to clean. | | |
| | `noisy_scale` | `2.5` | PDG scale at the first noisy schedule point. | | |
| | `clean_scale` | `2.5` | PDG scale at the final clean schedule point. | | |
| | `power` | `3.0` | Positive exponent used by the power curve. | | |
| | `start_step` | `0` | First active solver update, inclusive. | | |
| | `stop_step` | `None` | Last active solver update, inclusive. `None` selects the final update. | | |
| Constant PDG requires equal `noisy_scale` and `clean_scale`. Disabled PDG | |
| requires `enabled=False` and `mode=PdgMode.NONE`. | |
| Increasing the PDG scale sharpens the generated distribution. Moderate values | |
| improve image structure and fine detail at the cost of reduced variety. | |
| Pushing the scale too far can introduce structural defects, excessive | |
| contrast, and oversaturation. | |
| ### PDG curves | |
| | Enum | Value | Scale behavior | | |
| | --- | --- | --- | | |
| | `PdgCurve.CONSTANT` | `"constant"` | Uses one scale throughout inference. | | |
| | `PdgCurve.LINEAR` | `"linear"` | Interpolates linearly from `noisy_scale` to `clean_scale`. | | |
| | `PdgCurve.POWER` | `"power"` | Interpolates with `position ** power`. | | |
| ### PDG modes | |
| | Enum | Value | Behavior | | |
| | --- | --- | --- | | |
| | `PdgMode.NONE` | `"none"` | No PDG branch. Required when PDG is disabled. | | |
| | `PdgMode.FULL` | `"full"` | Guides from the middle-skipped path toward the full main path. | | |
| | `PdgMode.FULL_CONTRASTIVE` | `"full_contrastive"` | Uses negative text on the middle-skipped path, or learned unconditional text when no negative prompt is supplied. | | |
| | `PdgMode.THREE_QUARTER` | `"three_quarter"` | Guides from the 75 percent SPRINT path toward the full main path. | | |
| | `PdgMode.ALTERNATE_PDG_FIRST` | `"alternate_pdg_first"` | Alternates PDG on even updates and CFG on odd updates within the PDG window. | | |
| | `PdgMode.ALTERNATE_CFG_FIRST` | `"alternate_cfg_first"` | Alternates CFG on even updates and PDG on odd updates within the PDG window. | | |
| | `PdgMode.COMBINED_CFG_PDG` | `"combined_cfg_pdg"` | Evaluates CFG and PDG together and averages their guidance deltas. | | |
| | `PdgMode.PDG_WITH_ALTERNATING_CFG` | `"pdg_with_alternating_cfg"` | Applies PDG on every active update and adds CFG on odd updates. | | |
| | `PdgMode.CFG_TO_PDG` | `"cfg_to_pdg"` | Uses CFG before the PDG window, then uses full-path PDG inside the window. | | |
| The five compound modes require `cfg.scale`. `PdgMode.FULL_CONTRASTIVE` uses | |
| the PDG scale only. `start_step=0` with `PdgMode.CFG_TO_PDG` begins directly | |
| with PDG. | |
| `pdg_branch_conditioning(mode)` returns `PdgBranchConditioning.POSITIVE` or | |
| `PdgBranchConditioning.CONTRASTIVE` for the selected PDG alternative path. | |
| ## Self-attention gain | |
| `self_attention_gain` changes image self-attention only on the full main path. | |
| The model multiplies main-path attention queries by | |
| `exp(self_attention_gain)`. Negative values reduce the attention-logit scale | |
| and therefore increase the effective softmax temperature. This softens the | |
| main prediction and helps moderate PDG oversaturation. Text self-attention, | |
| cross-attention, and weak guidance paths retain their trained scales. | |
| The default is `-0.05`. A value of `0.0` uses the trained main-path | |
| self-attention scale without adjustment. Smaller images generally benefit from | |
| more negative values. As image size increases, the gain should move closer to | |
| zero. | |
| ## Custom configuration example | |
| ```python | |
| from canter import ( | |
| CanterInferenceConfig, | |
| CanterOutputType, | |
| CanterPipelineConfig, | |
| CfgGuidance, | |
| PdgCurve, | |
| PdgGuidance, | |
| PdgMode, | |
| Schedule, | |
| Solver, | |
| ) | |
| inference = CanterInferenceConfig( | |
| height=1024, | |
| width=1024, | |
| steps=40, | |
| solver=Solver.DPMPP_2M, | |
| schedule=Schedule.LINEAR, | |
| log_snr_shift=0.5, | |
| cfg=CfgGuidance( | |
| enabled=True, | |
| scale=3.0, | |
| start_step=0, | |
| stop_step=11, | |
| ), | |
| pdg=PdgGuidance( | |
| enabled=True, | |
| mode=PdgMode.CFG_TO_PDG, | |
| curve=PdgCurve.POWER, | |
| noisy_scale=2.0, | |
| clean_scale=2.5, | |
| power=3.0, | |
| start_step=12, | |
| stop_step=39, | |
| ), | |
| self_attention_gain=-0.03, | |
| sde_noise_multiplier=1.0, | |
| seed=123, | |
| generator=None, | |
| ) | |
| config = CanterPipelineConfig( | |
| inference=inference, | |
| output_type=CanterOutputType.PIL, | |
| ) | |
| image = pipe("A glass greenhouse during heavy rain", config=config).image | |
| ``` | |
| ## Prompt batches | |
| A sequence of prompts generates one image per prompt: | |
| ```python | |
| output = pipe( | |
| [ | |
| "A windswept beach under dark clouds", | |
| "A sunlit kitchen with white tiled walls", | |
| ] | |
| ) | |
| first, second = output.images | |
| ``` | |
| Batch size one is the primary inference path. The jagged backend packs active | |
| prompt tokens without padding them through text refinement and | |
| cross-attention. Prompts longer than 512 tokens are truncated with a warning. | |
| During inference, both positive and negative prompts accept | |
| `(text:weight)` emphasis. The markup is removed before tokenization and the | |
| non-negative numeric weight interpolates the matching tokens' projected keys | |
| and values from Canter's learned unconditional state. Prompt and unconditional | |
| keys are normalized separately before interpolation; there is no normalization | |
| afterward. For example, `(Compression artifacts:0.2)` weakens that phrase. | |
| Weight `1` is a no-op and weight `0` selects unconditional conditioning for the | |
| span. Parentheses without an explicit weight use ComfyUI's default `1.1`; | |
| negative weights and nested weighted groups are rejected. When one span covers | |
| the whole prompt, its weight also applies to the active special tokens. | |
| ## Latent generation | |
| Use `CanterInferenceEngine` to generate latents without loading or calling the | |
| VAE: | |
| ```python | |
| from canter import CanterComponents, CanterInferenceEngine | |
| components = CanterComponents.from_pretrained("data-archetype/canter") | |
| engine = CanterInferenceEngine(components) | |
| output = engine.generate("A mountain road in winter") | |
| latents = output.latents | |
| schedule = output.schedule | |
| ``` | |
| `CanterInferenceEngine.generate` accepts: | |
| | Parameter | Default | Description | | |
| | --- | --- | --- | | |
| | `prompts` | required | One string or a sequence of strings. | | |
| | `negative_prompts` | `None` | Optional negative prompt string or sequence for CFG and contrastive PDG branches. The count must match `prompts`; blank text uses learned unconditional conditioning. | | |
| | `config` | `CanterInferenceConfig()` | Latent inference settings. | | |
| | `initial_noise` | `None` | Optional float32 noise with shape `[batch, 128, height / 16, width / 16]`. | | |
| | `progress` | `None` | Optional callback receiving `(completed_updates, total_updates)`. | | |
| | `state_callback` | `None` | Optional synchronous callback receiving `(state, completed_updates, total_updates)` after every solver update. Prefer pipeline previews when full CUDA latent access is not required. | | |
| An empty or whitespace-only prompt uses the learned unconditional token without | |
| running the text encoder. Every prompt in a batch must be either blank or | |
| nonblank. | |
| The returned latents use float32 and channels-last memory format. | |
| ## Direct VAE encoding and decoding | |
| `CanterVae` exposes both directions of the pinned DINAC-AE latent interface: | |
| ```python | |
| import torch | |
| from canter import CanterVae | |
| vae = CanterVae.from_pretrained( | |
| "data-archetype/dinac_ae_d2", | |
| device="cuda", | |
| ) | |
| images = torch.zeros((1, 3, 1216, 832), device="cuda", dtype=torch.float32) | |
| latents = vae.encode(images) | |
| reconstructed = vae.decode(latents, height=1216, width=832) | |
| ``` | |
| Encoder inputs and decoder outputs are float32 BCHW tensors in the DINAC image | |
| range `[-1, 1]`. Encoded latents are whitened float32 tensors with 128 channels | |
| and a spatial stride of 16. | |
| ## Conditioning, guidance, and solver APIs | |
| The public inference API exposes text conditioning, guidance, schedules, | |
| solvers, and the latent-RGB preview projection as separate components: | |
| ```python | |
| import torch | |
| from canter import ( | |
| LOGSNR_SOLVER_START_EPS, | |
| CanterGuidanceConfig, | |
| CanterGuidedVelocity, | |
| LatentRgbProjection, | |
| Schedule, | |
| Solver, | |
| build_schedule, | |
| solve, | |
| ) | |
| conditional = engine.prepare_conditioning("A mountain road in winter") | |
| unconditional = engine.prepare_unconditional(batch=1) | |
| guidance = CanterGuidanceConfig() | |
| generator = torch.Generator(device=engine.device).manual_seed(42) | |
| velocity = CanterGuidedVelocity( | |
| components.model, | |
| conditional, | |
| unconditional, | |
| guidance, | |
| generator, | |
| ) | |
| schedule = build_schedule( | |
| Schedule.BETA, | |
| steps=guidance.steps, | |
| log_snr_shift=0.0, | |
| finite_noisy_endpoint_epsilon=None, | |
| device=engine.device, | |
| ) | |
| latents = solve( | |
| velocity, | |
| initial_noise, | |
| schedule, | |
| solver=Solver.ABM2, | |
| generator=generator, | |
| ) | |
| preview_projection = LatentRgbProjection.bundled() | |
| ``` | |
| `prepare_conditioning` validates a string or prompt sequence and performs the | |
| bundled text-encoder/refinement pass. An entirely blank batch returns learned | |
| unconditional conditioning; mixed blank and nonblank batches are rejected. | |
| `prepare_unconditional` creates learned conditioning for an explicit positive | |
| batch size. | |
| `CanterGuidanceConfig` contains the solver-step guidance state: step count, | |
| CFG, PDG, and self-attention gain. `CanterGuidedVelocity`, `build_schedule`, | |
| and `solve` preserve the same PDG, CFG, self-attention gain, stochastic | |
| generator, and float32 solver behavior as `CanterInferenceEngine.generate`. | |
| For `Solver.DPMPP_2M`, `Solver.DPMPP_2M_SDE`, `Solver.ER_SDE`, and | |
| `Solver.ER_SDE_BT`, pass `LOGSNR_SOLVER_START_EPS` as | |
| `finite_noisy_endpoint_epsilon`; other solvers use `None`. | |
| `LatentRgbProjection.bundled()` returns the validated CPU float32 1×1 projection | |
| weight and bias used by native previews. | |
| ## Pipeline metadata | |
| `pipe.metadata` records the installed Canter code version, resolved checkpoint | |
| release, weight dtype, source digests, text-encoder revision, VAE repository, | |
| and resolved immutable VAE revision. Applications that require reproducibility | |
| should store this metadata and pin both the package version and checkpoint tag. | |