Computer Vision 52

High-quality image to 3D model conversion for objects and human bodies.
SAM 3D is an online tool powered by Meta’s SAM 3D research models, designed to convert a single image into a high-quality 3D model of objects or human bodies in seconds. It generates high-fidelity 3D models from any photo, offering accurate object and human-body reconstruction with exceptional speed directly in the browser, without requiring setup or a GPU. The platform features two specialized architectures: SAM 3D Objects for scene-aware reconstruction of detailed 3D shapes, textures, and layouts from a single image, and SAM 3D Body for precise human digitization, generating accurate 3D human pose and shape estimations using the Meta Momentum Human Rig (MHR). It leverages a 'Human-in-the-Loop' data engine, trained on nearly 1 million physical world images, to achieve unmatched robustness in diverse real-world scenarios. SAM 3D supports standard 3D export formats and is fully open-source under the Apache 2.0 License.

Free motion capture software with AI-powered tools and a store of 3D assets.
Movmi is a free motion capture software designed to simplify human animation for animators and game developers. It supports various lifestyle locations and camera devices, utilizing advanced computer vision algorithms to capture face and body movements. Movmi also features a store with free 3D characters and animations, allowing users to apply captured motion to these characters. It offers AI-powered tools for generating poses from text and creating videos with AI-generated backgrounds.

AI-powered tool to remove Sora2 watermarks from videos.
SORA2WatermarkRemover is an advanced AI-powered tool specifically designed to remove Sora2 watermarks from videos. It leverages cutting-edge computer vision technology to intelligently detect and remove Sora2 branding, ensuring clean, professional results while preserving the original video quality. The tool supports batch processing, allowing users to remove AI watermarks from multiple videos simultaneously without requiring a login. It offers ultra-fast processing, no quality loss, and compatibility with various video formats, catering to content creators and video professionals.

AI data factory for building, operating, and staffing AI data.
Labelbox is a comprehensive AI data factory that provides software and services to help AI teams build, operate, and staff their data operations. It is designed to generate high-quality training data at scale for any AI project and to evaluate AI model performance. Labelbox offers a complete set of data solutions for modern AI and ML development, serving companies from startups to Fortune 500s.

Ultralytics provides vision AI tools and platforms for creating, training, and deploying ML models.
Ultralytics is a company focused on empowering individuals and businesses by providing vision AI tools. Their flagship product, Ultralytics HUB, is an AI platform designed for creating, training, and deploying machine learning models with a no-code interface. They also offer Ultralytics YOLO, a state-of-the-art AI tool for image classification, object detection, and instance segmentation. Ultralytics aims to simplify the process of building and deploying AI models, making it accessible to users regardless of their field of work.

Platform for streamlining AI data annotation and evaluation workflows.
SuperAnnotate is a comprehensive platform designed to streamline AI data workflows. It enables users to build feedback-driven annotation and evaluation pipelines for creating and managing high-quality AI data faster across infinite use cases. The platform centralizes all AI data work, supporting various data types including multimodal, image, video, NLP, and audio. It is built for cutting-edge AI initiatives such as RLHF, SFT, Agents, RAG, and general model evaluation, adapting to diverse workflows. SuperAnnotate integrates directly with existing AI stacks, data sources, and model training pipelines to reduce infrastructure complexities and facilitate modern AI development.

Appen provides data and services to improve AI model performance and accelerate AI development.
Appen provides high-quality, scalable data for AI models and applications. They offer an end-to-end platform, flexible services, and deep expertise to ensure the delivery of diverse data crucial for building foundation models and enterprise-ready AI applications. Appen supports the AI lifecycle by providing software to collect, curate, fine-tune, and monitor traditionally human-driven tasks, creating efficiencies through a trustworthy process. They offer AI training data, data annotation, data collection, LLM training data & services, multilingual AI, evaluation & benchmarking, supervised fine tuning, and off-the-shelf datasets.

Breakthrough AI processors for high-performance deep learning on edge devices.
Hailo offers breakthrough AI processors uniquely designed to enable high-performance deep learning applications on edge devices. Their processors are geared towards the new era of generative AI on the edge, in parallel to enabling perception and video enhancement through a wide range of AI accelerators and vision processors. Hailo aims to make high-performance AI widely available and affordable, helping make people’s lives safer, more convenient, and more productive, without compromising their privacy and security.

Explore experimental AI demos and research from Meta.
Meta AI Demos is a platform designed to showcase and allow users to try experimental demos featuring the latest AI research from Meta. It provides early access to cutting-edge AI tools developed by FAIR and other research teams across Meta, blending advanced research with creativity and technology. Users can explore various featured experiments and technical demos, test their functionalities, and contribute to the development of AI technologies that may be integrated into Meta's future products.

Reka is an agentic multimodal AI platform for visual understanding and data insights.
Reka is an AI research and product company that develops multimodal, modular intelligence solutions. Its platform, Reka Vision, specializes in agentic visual understanding and search across video, image, audio, and text, transforming raw unstructured data into deep insights and actions. Reka delivers complete AI solutions, from visual intelligence platforms for video editing and search to state-of-the-art web agents for researching complex questions, all powered by novel multimodal transformers built from scratch.

A digital media outlet providing the latest AI news, technologies, and applications.
Just AI News is a media outlet providing the latest artificial intelligence news. It offers up-to-date information on AI technologies, company developments, and real-world applications. The site is categorized into sections like Applications, Technologies, and Industries, making it easy to explore the AI world. Content is updated from Monday to Friday, with 6-10 news posts published daily. A curated newsletter is sent out every Tuesday.

AI custom software development and consulting, specializing in LLMs, MLOps, and computer vision.
deepsense.ai specializes in AI custom software development and AI consulting. They offer enterprise AI solutions, focusing on LLMs, MLOps, computer vision, and AI-powered automation to drive business growth. They provide guidance, strategy, and implementation to help businesses unlock the full potential of AI.

AI/ML solutions provider specializing in face-related computer vision and custom AI development.
Visage Technologies specializes in AI/ML solutions optimized for performance and compliance, leveraging over 20 years of consultancy and engineering experience. They offer custom development services and a range of products including visage|SDK™, FaceTrack, FaceAnalysis, FaceRecognition, and makeup|SDK. Their expertise spans various use cases such as driver monitoring, virtual makeup, face filtering, and eyewear try-on. They build AI-driven products for leading companies, focusing on edge AI solutions optimized for high accuracy and low latency, providing full support from development to deployment.

AI & Machine Learning consulting for revenue growth through NLP and computer vision.
Width.ai is an AI and Machine Learning consulting company focused on increasing revenue for businesses. They specialize in natural language processing and computer vision systems, helping businesses understand their revenue streams better and building tools to make them more profitable. They offer services ranging from MVP builds to full enterprise product development, focusing on generative AI implementations.

Meshy is a 3D AI platform for generating 3D models from text or images.
Meshy is an AI platform that allows three-dimensional (3D) content creation from text data and two-dimensional images. Users can leverage Meshy's features to convert text into a textured rendering, develop compelling 3D models from textual prompts, and transform an image into a 3D model with minimal effort. It has an AI-powered texture feature that lets users apply textures to 3D models seamlessly. Meshy offers a variety of art styles, including realistic, cartoon, and sculpture formats. Its interface is simple and intuitive, which makes it user-friendly, enabling non-experts to utilize it effectively. Meshy supports multiple languages and offers an Application Programming Interface (API) that allows its integration into other applications to expand its functionality. Users also have the advantage of previewing their work on a browser immediately after completion. It allows users to export their 3D models in various formats like FBX, OBJ, STL, BLEND, and USDZ for easy compatibility with other software applications.

Crowdsourcing platform for AI training data and data management services.
clickworker is a crowdsourcing platform that provides AI training data and other data management services. It leverages a global network of over 7 million Clickworkers to generate, validate, and label data. Services include AI dataset creation, content editing, surveys, internet research, categorization, tagging, and more. clickworker caters to industries like AI & Data Science, Research, eCommerce, Retail, and Digital Marketing.

Build any real-world vision app in minutes. From idea to live vision app, in three steps. Start with an idea. Type it in plain language or drop in a clip. No datasets, no labelling, no model selection, just say what to watch for. Then watch it build itself. Viso Now connects the feed, recognizes the scene, writes the detection logic and wires the alerts, all agentic, on auto-pilot, in front of you. Then simply refine and ship live. You can tune any rule, threshold or output, then publish. Your vision app runs live and pushes events wherever your team already works.
Viso Now is a visual AI platform that lets anyone turn video footage into a working computer vision application in minutes. Simply upload a video and describe what you want to identify, monitor or understand, and Viso Now builds a vision agent that can analyse the scene, reason about what is happening and deliver useful insights. From detecting safety risks and quality issues to monitoring operations, traffic, healthcare environments or entirely new ideas, it makes visual intelligence accessible without data labelling, model training or machine learning expertise. It is free to start, with daily credits included.
Neuron Q provides AI solutions, including firearm detection system for enhanced safety.
Neuron Q is a leader in AI-powered solutions, offering cutting-edge technologies in Artificial Intelligence, Machine Learning, and Computer Vision to revolutionize industries and create a safer, more efficient future. Their AI-powered surveillance system, Project Argus, is designed to detect firearms and prevent potential threats. It integrates seamlessly with existing camera systems, offering real-time alerts and operates autonomously, reducing the need for constant human monitoring.

Unbound AI is an image generation studio empowering creators with AI and blockchain solutions.
Unbound AI is an intuitive platform that blends diffusion models trained on multiple image styles. It provides tools for creators and startups to generate high-quality images and graphic designs. It aims to empower creators with a complete Image Generation Studio, offering features like NFT minting, NSFW content generation (with KYC), and an open-source program. The platform also includes Unbound Vision, an AI Computer Vision solution for retail, Unbound Network, an open-source blockchain, and Unbound Identity, a customizable blockchain Identity Solution.

AI video intelligence platform for searching, analyzing, and generating text from video content.
TwelveLabs offers an AI-powered video intelligence platform that uses multimodal models (Marengo/Pegasus) to search, analyze, and generate text from video content at scale. It enables users to find anything, discover deep insights, analyze, remix, and automate workflows across their entire video content. TwelveLabs' AI surpasses benchmarks from cloud majors and open-source models, providing world-class accuracy and customization.

AI-powered visual search and analysis for teams.
CoreViz is a comprehensive visual AI platform designed to be a visual co-pilot for teams and organizations. It allows users to upload thousands of images and videos and then ask anything about them, functioning as an advanced Google Photos for professional use. At its fullest, CoreViz acts as a forensics expert, radiologist, and geospatial analyst, automating visual insights. It leverages advanced computer vision and machine learning technologies to automatically understand, index, and make visual data searchable using natural language queries, without requiring any coding expertise. The platform unlocks the power of images and videos through AI-powered search, analysis, and insights, trusted by over 100 organizations worldwide.

Robotics simulation platform for training AI without physical robots.
Lucky Robots offers a robotics simulation platform, a virtual training boot camp for robots. It utilizes cutting-edge technologies to rapidly iterate, train, and test robot models, abilities, and potential in a digital environment. The platform enables training end-to-end AI for robotics without physical robots by generating infinite synthetic data, building on existing models, and seamlessly iterating, training, and testing models. Lucky Robots aims to make robotics accessible to regular software engineers by decoupling it from ROS and physical hardware, making it usable with natural language.

Image Recognition API for tagging, categorization, visual search, and content moderation.
Imagga Image Recognition API provides solutions for image tagging & categorization, visual search, content moderation. Available in the Cloud and On-Premise. Plans available to suit all needs. Subscription plans based on API calls.
Fyusion offers AI-powered vehicle damage detection and 3D imaging solutions.
Fyusion provides vehicle damage detection and 3D vehicle imaging solutions using AI. Their technology offers comprehensive condition reporting, interactive 3D imaging, and AI-based collision reports for the automotive industry.