Category · 79 repos
Computer Vision
Detection, segmentation, tracking, pose, depth and vision-language models. Ranked by star velocity over the last 24 hours.
Showing 1–50 of 79
Official PyTorch implementation of VisionHOPE: Visual Backbones as Self-Modifying Learning Systems.
π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.
Eyes are All JEV Needs - real time visual devisions from RGB, video and RGB-D cameras.
🚀🚀 A multimodal System One decision model that gives calibrated answers to typed questions about screens, photos, video and text in one forward pass.
A skill that turns SRT subtitles into flowing whiteboard handwriting animation on warm beige paper: mask-region choreography + continuous stream strokes (ink→color).
Train a Jev-like multimodal model by yourself. System One Model, now with vision.
Multi-tenant AI SaaS platform — autonomous agentic routing (RAG/CNN/chat), PyTorch vision, FAISS retrieval, Celery/Redis pipelines, Stripe billing, and full observability.
vLLM Omni backend plugin for diffusion and multimodal generation on AWS Trainium
An offline-first photo manager for large local libraries
Simple OpenCV based face smoothing
Cross-platform library for converting natural raster images into clean SVGs - fast, precise and configurable.
Fast and Accurate ML in 3 Lines of Code
One of the Seetaface face recognition engine module
One of the Seetaface face recognition engine module
YOLO-OMNI: A cross-domain real-time object detection framework for fisheye, drone, panorama, game-to-real, and mixed-camera scenarios.
An autonomous, context-aware AI desktop companion. Built with Python, featuring real-time screen vision, custom ONNX voice synthesis, active window tracking, and a dynamic floating UI with reactive facial expressions.
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
AI-powered agents for automating 3D content workflows using Vision-Language Models (VLMs). Content Agents analyze 3D assets and automate material assignment, physics property classification, and texture generation for USD files.
Dealing with all unstructured data, such as reverse image search, audio search, molecular search, video analysis, question and answer systems, NLP, etc.
Run Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server. 6 fighters incl. Muse-Glimmer 30B (now multimodal — reads images, abliterated), Gemma 4 31B, Qwen 3.5 122B (65 tok/s), DeepSeek V4 Flash (1M ctx). Private, offline, airgap-ready. Built for NDA / legal / h
[EMNLP-2024] Build multimodal language agents for fast prototype and production
Halo is an open-source framework built by White Circle for training large language and multimodal models
Virtual Tryon For complex tradition outfits
https://arxiv.org/abs/2606.28344. The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/
Build a Raspberry Pi wildlife camera step by step — from Python image capture and motion detection to edge AI.
Visualize, query, and stream to train on multimodal robotics data.
Metric depth from one RGB frame and a phone LiDAR: LiDAR-prompted MoGe-3 with real-time on-device models (Core ML + Metal)
[ICCV 2025] LayerD: Decomposing Raster Graphic Designs into Layers
Eagle: Frontier Vision-Language Models with Data-Centric Strategies
The SDK for Jetpac's iOS Deep Belief image recognition framework
Make any agent harness multimodal-native.
3,516 curated Figure 1 / teasers from ICLR, ICML, NeurIPS, CVPR, ACL, AAAI (2023-2026), with tier badges and FigureForge retrieval-augmented figure drafting
Remove visible and invisible AI watermarks and provenance metadata from images and video. Python library and CLI for SynthID, C2PA, EXIF, IPTC, XMP, and common generative-AI marks.
This repository collects and organises state‑of‑the‑art papers on spatial reasoning for Multimodal Vision–Language Models (MVLMs).
Real-event-anchored benchmark for detecting AI-generated videos in real-world crisis settings.
A real-time Python-based ADAS pipeline that combines computer vision, YOLOv8 object detection, and sensor fusion to deliver lane detection, collision prediction, traffic sign recognition, and live driver guidance through an interactive HUD.
Open-source iOS app connecting Meta Ray-Ban smart glasses to AI — 5 backends (on-device MLX models, Apple Intelligence, OpenAI, Gemini Live, OpenClaw), on-device neural voice, face recognition & live web search. Private and offline-capable.
An open source platform for visual-inertial navigation research.
A complete, structured hub for learning Artificial Intelligence — covering AI, Machine Learning, Deep Learning, and Data Science with books, roadmaps, and curated resources from beginner to advanced.
Learning OpenCV 4 Computer Vision with Python 3 – Third Edition, published by Packt
[ICCV2021] Learning to Track Objects from Unlabeled Videos
Generative AI satellite cloud removal system using SAR-guided diffusion bridges for ISRO LISS-IV imagery — ISRO Bharatiya Antariksh Hackathon 2026 Grand Finale Finalist
Manga translation app powered by AI
Converseen is a batch image converter and resizer
FAST-LIVO2: Fast, Direct LiDAR-Inertial-Visual Odometry
Fun real-world computer vision demos!
An IoT-based multi-person fall detection system using YOLO and ESP32-CAM. This project integrates advanced object detection with ESP32-CAM for real-time fall detection in multi-person environments, suitable for smart surveillance, healthcare, and security systems
A ComfyUI custom node designed for advanced image background removal and object, face, clothes, and fashion segmentation, utilizing multiple models including RMBG-2.0, INSPYRENET, BEN, BEN2, BiRefNet, SDMatte, SAM, SAM2, SAM3 and GroundingDINO.
Desktop agent built on H Company's Holo3 vision-language models
Tool for processing film negatives.