Skip to main content

Category · 79 repos

Computer Vision

Detection, segmentation, tracking, pose, depth and vision-language models. Ranked by star velocity over the last 24 hours.

Showing 1–50 of 79

1
PSRben/VisionHOPE

Official PyTorch implementation of VisionHOPE: Visual Backbones as Self-Modifying Learning Systems.

PythonComputer Vision
2
ruvnet/RuView

π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.

RustComputer VisionSelf-Hosted Apps
3
CharlesFeng0314/JEV_sees

Eyes are All JEV Needs - real time visual devisions from RGB, video and RGB-D cameras.

PythonComputer VisionRobotics & Embodied AI
4
OmniJev/OneJev

🚀🚀 A multimodal System One decision model that gives calibrated answers to typed questions about screens, photos, video and text in one forward pass.

PythonComputer VisionBrowser Automation & Scraping
5
geeklee/srt-whiteboard-animation

A skill that turns SRT subtitles into flowing whiteboard handwriting animation on warm beige paper: mask-region choreography + continuous stream strokes (ink→color).

PythonComputer VisionProductivity & Notes
6
Liuziyu77/Valen

Train a Jev-like multimodal model by yourself. System One Model, now with vision.

PythonComputer Vision
7
Nikhil-creat/omniscale-enterprise-pro

Multi-tenant AI SaaS platform — autonomous agentic routing (RAG/CNN/chat), PyTorch vision, FAISS retrieval, Celery/Redis pipelines, Stripe billing, and full observability.

PythonComputer VisionRAG & Knowledge Bases
8
aws-neuron/vllm-omni-neuron

vLLM Omni backend plugin for diffusion and multimodal generation on AWS Trainium

PythonComputer VisionLLM Inference & Serving
9
julyx10/lap

An offline-first photo manager for large local libraries

VueComputer VisionSelf-Hosted Apps
10
yakhyo/face-smoothing

Simple OpenCV based face smoothing

PythonComputer Vision
11
Ryan-Millard/Img2Num

Cross-platform library for converting natural raster images into clean SVGs - fast, precise and configurable.

C++Computer Vision
12
autogluon/autogluon

Fast and Accurate ML in 3 Lines of Code

PythonComputer VisionMachine Learning & Data Science
13
linuxdeepin/seetaface-authorize

One of the Seetaface face recognition engine module

C++Computer Vision
14
seanwallawalla/LinuxDeepIn_SeetaFace-Anti-Spoofing-X

One of the Seetaface face recognition engine module

C++Computer Vision
15
z637826/yolo-omni

YOLO-OMNI: A cross-domain real-time object detection framework for fisheye, drone, panorama, game-to-real, and mixed-camera scenarios.

PythonComputer VisionRobotics & Embodied AI
16
AhmadHassan-BTed/B

An autonomous, context-aware AI desktop companion. Built with Python, featuring real-time screen vision, custom ONNX voice synthesis, active window tracking, and a dynamic floating UI with reactive facial expressions.

PythonComputer VisionText to Speech & Voice
17
Blaizzy/mlx-vlm

MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.

PythonComputer VisionLocal LLMs
18
NVIDIA-Omniverse/usd-content-agents

AI-powered agents for automating 3D content workflows using Vision-Language Models (VLMs). Content Agents analyze 3D assets and automate material assignment, physics property classification, and texture generation for USD files.

PythonComputer VisionWorkflow Automation & No-Code
19
milvus-io/bootcamp

Dealing with all unstructured data, such as reverse image search, audio search, molecular search, video analysis, question and answer systems, NLP, etc.

Jupyter NotebookComputer VisionRAG & Knowledge Bases
20
nicedreamzapp/claude-code-local

Run Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server. 6 fighters incl. Muse-Glimmer 30B (now multimodal — reads images, abliterated), Gemma 4 31B, Qwen 3.5 122B (65 tok/s), DeepSeek V4 Flash (1M ctx). Private, offline, airgap-ready. Built for NDA / legal / h

PythonComputer VisionLocal LLMs
21
om-ai-lab/OmAgent

[EMNLP-2024] Build multimodal language agents for fast prototype and production

PythonComputer VisionChatbots & LLM Apps
22
whitecircle/halo

Halo is an open-source framework built by White Circle for training large language and multimodal models

PythonComputer VisionFine-tuning & Training
23
Amrutha-dev25/TradViTon

Virtual Tryon For complex tradition outfits

PythonComputer Vision
24
StarTrail-org/PixelRAG

https://arxiv.org/abs/2606.28344. The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/

PythonComputer VisionRAG & Knowledge Bases
25
murraystokely/cubscout-trailcam

Build a Raspberry Pi wildlife camera step by step — from Python image capture and motion detection to edge AI.

PythonComputer VisionSelf-Hosted Apps
26
rerun-io/rerun

Visualize, query, and stream to train on multimodal robotics data.

RustComputer VisionRobotics & Embodied AI
27
sergmister/PromptMoGe

Metric depth from one RGB frame and a phone LiDAR: LiDAR-prompted MoGe-3 with real-time on-device models (Core ML + Metal)

PythonComputer VisionLocal LLMs
28
CyberAgentAILab/LayerD

[ICCV 2025] LayerD: Decomposing Raster Graphic Designs into Layers

PythonComputer Vision
29
NVlabs/Eagle

Eagle: Frontier Vision-Language Models with Data-Centric Strategies

PythonComputer Vision
30
ExpeditionEstates/recognition

The SDK for Jetpac's iOS Deep Belief image recognition framework

JavaScriptComputer VisionMobile & Desktop Apps
31
QwenLM/Qwen-MM-Plugins

Make any agent harness multimodal-native.

PythonComputer VisionAI Agents
32
qwdwqfwq/topconf-paper-figure-gallery

3,516 curated Figure 1 / teasers from ICLR, ICML, NeurIPS, CVPR, ACL, AAAI (2023-2026), with tier badges and FigureForge retrieval-augmented figure drafting

HTMLComputer VisionRAG & Knowledge Bases
33
wiltodelta/remove-ai-watermarks

Remove visible and invisible AI watermarks and provenance metadata from images and video. Python library and CLI for SynthID, C2PA, EXIF, IPTC, XMP, and common generative-AI marks.

PythonComputer VisionImage Generation & Editing
34
zhengxuJosh/Awesome-Multimodal-Spatial-Reasoning

This repository collects and organises state‑of‑the‑art papers on spatial reasoning for Multimodal Vision–Language Models (MVLMs).

Computer Vision
35
Shuo-Liang-0111/RA-Bench

Real-event-anchored benchmark for detecting AI-generated videos in real-world crisis settings.

PythonComputer VisionAI Safety & Red Teaming
36
SpoodermanCodes/adas-driver-assistance-system

A real-time Python-based ADAS pipeline that combines computer vision, YOLOv8 object detection, and sensor fusion to deliver lane detection, collision prediction, traffic sign recognition, and live driver guidance through an interactive HUD.

PythonComputer Vision
37
rayl15/OpenVision

Open-source iOS app connecting Meta Ray-Ban smart glasses to AI — 5 backends (on-device MLX models, Apple Intelligence, OpenAI, Gemini Live, OpenClaw), on-device neural voice, face recognition & live web search. Private and offline-capable.

SwiftComputer VisionLocal LLMs
38
rpng/open_vins

An open source platform for visual-inertial navigation research.

C++Computer VisionRobotics & Embodied AI
39
bishwaghimire/ai-learning-roadmaps

A complete, structured hub for learning Artificial Intelligence — covering AI, Machine Learning, Deep Learning, and Data Science with books, roadmaps, and curated resources from beginner to advanced.

Computer VisionMachine Learning & Data Science
40
PacktPublishing/Learning-OpenCV-4-Computer-Vision-with-Python-Third-Edition

Learning OpenCV 4 Computer Vision with Python 3 – Third Edition, published by Packt

PythonComputer Vision
41
VISION-SJTU/USOT

[ICCV2021] Learning to Track Objects from Unlabeled Videos

PythonComputer Vision
42
jaypatel29042008-glitch/LISS-IV-Cloud-Removal

Generative AI satellite cloud removal system using SAR-guided diffusion bridges for ISRO LISS-IV imagery — ISRO Bharatiya Antariksh Hackathon 2026 Grand Finale Finalist

PythonComputer VisionMachine Learning & Data Science
43
meangrinch/MangaTranslator

Manga translation app powered by AI

PythonComputer VisionNLP & Language Models
44
Faster3ck/Converseen

Converseen is a batch image converter and resizer

C++Computer Vision
45
hku-mars/FAST-LIVO2

FAST-LIVO2: Fast, Direct LiDAR-Inertial-Visual Odometry

C++Computer Vision3D, Avatars & Digital Humans
46
jeremyipark/vision-demos

Fun real-world computer vision demos!

PythonComputer Vision
47
tonlongthuat/Real-Time-Fall-Detection

An IoT-based multi-person fall detection system using YOLO and ESP32-CAM. This project integrates advanced object detection with ESP32-CAM for real-time fall detection in multi-person environments, suitable for smart surveillance, healthcare, and security systems

PythonComputer Vision
48
1038lab/ComfyUI-RMBG

A ComfyUI custom node designed for advanced image background removal and object, face, clothes, and fashion segmentation, utilizing multiple models including RMBG-2.0, INSPYRENET, BEN, BEN2, BiRefNet, SDMatte, SAM, SAM2, SAM3 and GroundingDINO.

PythonComputer VisionImage Generation & Editing
49
hcompai/holo-desktop-cli

Desktop agent built on H Company's Holo3 vision-language models

PythonComputer Vision
50
marcinz606/NegPy

Tool for processing film negatives.

PythonComputer Vision