Category · 79 repos
Computer Vision
Detection, segmentation, tracking, pose, depth and vision-language models. Ranked by star velocity over the last 24 hours.
Showing 51–79 of 79
Server-side video workflows for agents: ingest, understand, search, edit, stream.
Aliyun CAPTCHA 2.0 V3 (INPAINTING) automatic slider solver — jsdom environment spoofing + multimodal AI image recognition to locate the missing piece
RMT (RuoMengTu) is a free, open-source macro tool built on AHKv2. Let the code handle the tedious work—you have more meaningful things to do.
Privacy-first employee attendance verification with YOLO and encrypted face embeddings
Document scanning app
Otvoreni istraživačko-razvojni mobilni projekt za lokalnu analizu prometnih scena.
Computer Vision Annotation Tool (CVAT) is a leading platform for building high-quality visual datasets for vision AI. It offers open-source, cloud, and enterprise products, as well as labeling services, for image, video, and 3D annotation with AI-assisted labeling, quality assurance, team collaborat
HVPNet: A Bio-Inspired Network for General Salient and Camouflaged Object Detection(Neural Network 2026)
[CVPR 2026] Official release of "Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning"
Open source software for autonomous drones.
Synchronizes Lighthouse devices with SLAM headsets using a Follow SLAM HMD mode that keeps the headset pose authoritative and continuously aligns Lighthouse tracking to it.
Ultralytics YOLO27, YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
List of Computer Science courses with video lectures.
Deeplake is AI Data Runtime for Agents. It provides serverless postgres with a multimodal datalake, enabling scalable retrieval and training.
Streaming 3D reconstruction from only video input: Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction
A collection of CVPR 2026 papers and open-source projects.
The most lightweight image viewer for Windows, supporting a wide range of image formats.
Hardware development resource library: MC520 encoder motor, TB6612 voltage-regulator driver, RT1064 core board and mainboard, 0.96-inch OLED, and OpenART visual image transmission.
12 Weeks, 24 Lessons, AI for All!
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as N
Open-source vacuum robot cleaner
Pure C# PP-OCRv6 inference library: hand-written kernels, GEMM, and other operators for multiple platforms, with very high performance, low memory requirements, and high accuracy. Includes a managed ONNX interpreter and does not depend on Paddle Inference, ONNX Runtime, or native OpenCV libraries.
NVR with realtime local object detection for IP cameras
Interactive deep learning book with multi-framework code, math, and discussions. Adopted at 500 universities from 70 countries including Stanford, MIT, Harvard, and Cambridge.
An automated Electronic Know-Your-Customer (E-KYC) system designed to verify user identities securely. This project combines Optical Character Recognition (OCR) for document text extraction and Biometric Facial Recognition with Liveness Detection to prevent identity fraud.
CNN for EuroSAT satellite image classification, focused on checkpointing and reloading model weights with TensorFlow 2 / Keras.
An extension of Open3D to address 3D Machine Learning tasks
Browser & Web Worker focussed image codec wasm bundles derived from the Squoosh App.
An end-to-end AI system for brain tumor classification and MRI tumor segmentation, built with PyTorch, CNNs, U-Net, and transfer learning.