🌳 Visual Tree
Upload an image → get its condensed tree (hover a node), the reveal animation, and a reveal video (super-resolution) — all best-first by CLS attention.
Encoder
Compare encoder families: DINOv3 (self-supervised, semantic), MAE (pixel-reconstruction — more texture/appearance driven), CLIP (language-aligned — 16×16 grid, patch-14), DreamSim (tuned on human similarity judgments).
ViT-H+/7B are large — slow on the free CPU; ViT-7B realistically needs a GPU Space. First use of each model downloads its weights (one-time; DreamSim is ~1 GB).