Projects based on SigLIP (Zhai et. al, 2023) and Hugging Face transformers integration 🤗
-
Updated
Feb 21, 2025 - Jupyter Notebook
Projects based on SigLIP (Zhai et. al, 2023) and Hugging Face transformers integration 🤗
Inference and fine-tuning examples for vision models from 🤗 Transformers
[ICCVW 2025] LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
[CVPR'25-Demo] Official repository of "TryOffDiff: Virtual-Try-Off via High-Fidelity Garment Reconstruction using Diffusion Models".
[NeurIPS 2024] AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation
Local-first cinematic visual archive for filmmakers — search stills by look, mood & technique; ingest video/URLs; moodboards. Own your frames (FilmGrab / Flim / Kive style, on your machine).
本项目以应用为主出发,结合了从基础的机器学习、深度学习到目标检测以及目前最新的大模型,采用目前成熟的 第三方库、开源预训练模型以及相关论文的最新技术,目的是记录学习的过程同时也进行分享以供更多人可以直接进行使用。
[ICLR 2026] The implementation of the paper Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
[ICLR 2025] - Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
Fine-Tuning SigLIP 2 for Single/Multi-Label Image Classification. Image classification vision-language encoder model fine-tuned for Image Classification Tasks
Low-latency ONNX and TensorRT based zero-shot classification and detection with contrastive language-image pre-training based prompts
Download flickr8k, flickr30k image caption datasets
Official PyTorch implementation of the WACV 2025 Oral paper "Composed Image Retrieval for Training-FREE DOMain Conversion".
A minimal, but effective implementation of CLIP (Contrastive Language-Image Pretraining) in PyTorch
MODA: open fashion retrieval benchmark and models by Hopit AI. MODA (203M, open source), MODA Pro Lite (213M, open weights), MODA Pro (hosted). Full-corpus benchmarks vs FashionSigLIP, SigLIP-SO400M and ZooClaw — one harness, losses shown. #1 open model on LookBench.
Spotlight-style local search for everything you saved and forgot: GitHub stars, local files, images, and bookmarks. Privacy-first, no full-disk scanning, fully on-device.
Code for Post-hoc Probabilistic Vision-Language Models
Chitrarth: Bridging Vision and Language for a Billion People
Meme search and discovery engine using OpenAI CLIP and Salesforce BLIP
An AI powered Video Serach Engine with google's SigLIP and Qdrant. It allows to search objects or key moments in videos just using natural language.
Add a description, image, and links to the siglip topic page so that developers can more easily learn about it.
To associate your repository with the siglip topic, visit your repo's landing page and select "manage topics."