PromptAI

Vision — AI news & research

Image, video and multimodal AI — diffusion models, generation, segmentation and visual understanding. Updated continuously from 20+ curated sources.

AgentsVisionRoboticsPolicyResearchToolsLLMs
hn · Vision

Show HN: A working 3D model of an Enigma machine

I watched the excellent Veritasium video [1] on the Enigma machine, and watched the full animation by Jared Owen [2], but was still a bit confused on how the inner mechanics of an Enigma machine work…

hn · Vision

Anthropic's IPO prospectus shows AI vision, surging costs

marktechpost · Vision

A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation

A comprehensive coding tutorial on Google Research's Massive Sound Embedding Benchmark (MSEB), demonstrating how to implement custom sound encoders, drive classification, clustering, retrieval, and segmentation evaluato…

hn · Vision

An Antidote to Roko's Basilisk

I have a theory: a mirror image of Roko’s Basilisk. It’s a future AI, millions of years in the future, that values the complexity of every human life so deeply that, after it and humanity have master…

hn · Vision

SNL Weekend Update: Anthropic CEO Dario Amodei on A.I.'S Threat to Humanity [video]

hn · Vision

Turning GLM-5.3-Flash into a Jev-like decision model

We found an approach to get Jev-like properties from standard LLMs like GLM-5.3-Flash. The core idea is to craft the input prompt so that the first output token answers the question. This makes it po…

hn · Vision

Show HN: A Claude Code skill to analyze your chess games

Hello HN, It started as an experiment: can Claude play chess properly if it uses vision instead of PGN notation? Somehow it can. The next experiment was to see whether Claude + Stockfish could explai…

hn · Vision

Alan Kay: Shannon gave us a way of dealing with noisy channels [video]

The following explanation is taken from https://news.ycombinator.com/item?id=49622607 : Alan Kay performs an improvisational avant garde layered audio feedback loop about Claude Shannon, live online…

hf · Vision

Accelerating vision-language models with LFM2.5-VL-DSpark

analytics · Vision

How AI Is Reshaping Corporate Video Production for Tech Brands

hn · Vision

Meta takes down a critical video about meta AI Glasses after filming at Meta

hn · Vision

Show HN: InstinctFlash – Run 5B world-action models in real time on Jetson Thor

Hey HN, Guanming here, cofounder of General Instinct. We just released InstinctFlash, a high-performance serving framework for robotics models. It’s licensed under AGPL-3.0. On Jetson Thor, we see sp…

openai · Vision

Higgsfield AI ships new video features in a day with GPT-6 Astra

With GPT-6 Astra, Higgsfield AI makes video ad creation easier for small businesses and brings new creative tools to market faster.

hn · Vision

Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

Hey HN, Toby from Nari Labs here. We've been working on making OSS speech models super-fast. Last year, we built Dia, the first OSS text-to-speech model capable of doing natural dialogue. Since then,…

hn · Vision

Show HN: MultiMatte, a Promptable Image Background Removal Model

Hey HN, I'm Shreyash from Feyn. We help companies build custom models from their data. Today we're releasing MultiMatte, a background removal model you can aim with words. Name an object in your imag…

hn · Vision

The paradox of diffusion distillation (2024)

hf · Vision

NeoMME: an efficient Multimodal-native and Multilingual Encoder

hn · Vision

Launch HN: RonanRX (YC S26) – Personalized Peptides and GLP-1s

Hi HN, I’m Lloyd, one of two founders of RonanRx ( https://ronanrx.com/ ). We are building a vertically integrated pharmaceutical company with software for prescribing, telehealth, compounding, manuf…

hn · Vision

Launch HN: RonanRX (YC S26) – Personalized Peptides and GLP-1s

Hi HN, I’m Lloyd, one of two founders of RonanRx. We are building a vertically integrated pharmaceutical company with software for prescribing, telehealth, compounding, manufacturing, and delivery. W…

hn · Vision

Launch HN: Nori Robotics (YC S26) – A low-cost humanoid robot for development

Hey HN, I’m Antonio from Nori Robotics ( https://norirobotics.com ). We build a $1,688 bimanual mobile robot in San Francisco for robotics developers and researchers. I started working on Nori while…

tc · Vision

Clipto uses AI to search terabytes of video and is now valued at $250M

The three-year-old startup says it reached $15 million in ARR and profitability before raising its latest $15 million round.

hn · Vision

Launch HN: Hebbian Robotics (YC S26) – Build scalable robotics data pipelines

Hi HN, we’re Brandon and Kingston, the founders of Hebbian Robotics. We built HFlow ( https://github.com/Hebbian-Robotics/hflow ), an SDK that turns multimodal recordings from robots and human operat…

hn · Vision

How to build a diffusion language model

hn · Vision

Show HN: Academa – Long-form STEM lecture videos generated by LLMs

Hi HN, we are Sina Atalay and Abdullah Geduk, co-founders of Academa. We are both PhD students. We thought: what if lecture videos were written as code and compiled into video using computer graphics…

hn · Vision

Continuous Diffusion Language Models (CDLM's)

hn · Vision

Laion Big Video Dataset

tc · Vision

Stability AI, maker of image generator Stable Diffusion, raises $76 million in fresh funding

The company's new fundraising total now stands at $232 million.

hn · Vision

Show HN: Screen memory without screenshots, just text to Markdown

It's a macOS menu bar app that reads the text of your focused window every few seconds through the Accessibility API. No screenshots, no video, or OCR. It writes plain markdown, one file per day, int…

analytics · Vision

Apple Job Cuts Hit Siri, Vision Pro Teams During Major AI Restructuring

hn · Vision

GPT 5.6 Sol is the best "vision" model OpenAI ever released

tc · Vision

Why people aren’t buying Mark Zuckerberg’s AI future

On the latest episode of Equity podcast, we discuss why not everyone is buying Zuckerberg’s vision.

hn · Vision

Show HN: Live Claude Usage HUD for a $38 Thermalright Trofeo Vision LCD

hn · Vision

Show HN: OpenCode Senses, An insanely fast and highly accurate vision plugin

The vision plugin for OpenCode that truly understands images. Inspect, read, and reason about any screenshot or picture with deeper understanding than any other plugin — fully local, private, and fre…

hf · Vision

LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge

hn · Vision

Launch HN: Discovered Materials (YC P26) – AI agents to discover new materials

Hey HN, we're Advaith and Akash from Discovered Materials ( https://discoveredmaterials.com/ ). We build AI agents that discover new materials for the semiconductor industry. GPUs today have a heat p…

tc · Vision

General Catalyst leads $1.1B round into 2-month-old River AI

River AI, a startup founded by xAI co-founder Igor Babuschkin, has a fascinating vision for personal agents and secured $1.1 billion out of the gate.

hn · Vision

Launch HN: Keet (YC S24) – An app to create video courses on anything

Hi HN! We’re Zack and Tommy the Co-Founders of Keet ( https://trykeet.com ). We are building a mobile app that generates courses on any topic, with short videos for explanation and games for reinforc…

tc · Vision

Meta’s new Glimmer AI model offers a hint at Zuckerberg’s personal intelligence vision

Meta’s new open-weight Muse Glimmer model offers a glimpse of Mark Zuckerberg’s personal superintelligence vision, as well as the emerging divide between AI users can own and access.

hn · Vision

Google Caught AI Faking Creativity in Every Office in America [video]

marktechpost · Vision

Mistral AI Releases Shieldstral 1.0 3B: An Open-Weights Policy-Adaptive Multimodal Safety Classifier Matching Models 7× Its Size

Mistral AI has released Shieldstral 1.0 3B, an open-weights, policy-adaptive multimodal safety classifier that frames content moderation as a single yes/no question instead of a fixed harm taxonomy. Operators supply the…

marktechpost · Vision

Building a Multimodal RAG Pipeline with NVIDIA NeMo Retriever, Hosted NIMs, LanceDB, Reranking, and Grounded Generation

In this tutorial, we build an advanced multimodal retrieval-augmented generation pipeline with NVIDIA NeMo Retriever. We begin by configuring a Python 3.12 environment, installing the required packages, and performing o…

hn · Vision

Mistral's Shieldstral: 3B open-weights model for multimodal moderation

hn · Vision

Show HN: Kedge – Full-stack cloud with forkable VM snapshots and global SQLite

I'm building Kedge, a globally distributed platform for stateful serverless apps. Here's how you make a simple static site: `echo '# Hello world!' | ssh kedge.dev' I helped build Fly.io for 4 years a…

hn · Vision

Robotics development made dead simple (open source)

Hey everyone, My team and I have been working hard on this project: https://peppy.bot It's a direct replacement for ROS 2. We already have the OpenArm robot (https://openarm.dev) working on the platf…

hn · Vision

Show HN: FeyNoBg – Automatic background removal model and training library

Hey HN, I’m Shreyash from Feyn. We help companies build custom models from their data. Today, we’re releasing FeyNoBg, an automatic background removal model. Alongside it, we're open-sourcing NoBg, t…

hn · Vision

D-FINE-seg – detection, instance and semantic segmentation in one model

tc · Vision

Midjourney acquired the astrology app Co-Star

The AI lab Midjourney continues to expand its purview beyond image and video generation.

hf · Vision

Bringing Nunchaku 4-bit Diffusion Inference to Diffusers

Browse briefing issues →

Get the briefing

The one story that matters, 5 headlines and the paper everyone's citing — every Tuesday, free.

Subscribe free