Images, Audio & Video
50 items tagged with this topic
Recent
DeepSeek-V4-Flash-Vision-Exp Release: Multimodal API Now Live | DeepSeek API Docs
DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀
Older
From Restoring Sight to Reimagining the Brain, with Max Hodak
Speaker 1 | 00:00 - 00:20 The brain very literally, very clearly, plainly is a computer. You can solve computational problems by arranging matter in a certain way and then like taking your hands off and pressing go. We talk about being a b…
China's Best Music in 2026 (So Far)
Jake Newby is the author of the Concrete Avalanche Substack, which covers the best music coming out of China.
The DeepSeek Thesis
a very different kind of guy
Most inspirational video you’ll watch today https://t.co/9ba2DwPMeZ
Most inspirational video you’ll watch today https://t.co/9ba2DwPMeZ
all of the recent proc gen art, video editing and 3d game demos recently have made me update to…
all of the recent proc gen art, video editing and 3d game demos recently have made me update towards LLM coding models being better at a lot of creative work than diffusion models
It's free to try and works with your existing Claude Code or Codex subscription right now. Just…
It's free to try and works with your existing Claude Code or Codex subscription right now. Just make a new directory and start Codex or CC (on command line OR on Desktop, slightly better on Desktop) Just paste this imag…
My hottest take is that brand marketing is going to be THE major differentiator and one of the…
My hottest take is that brand marketing is going to be THE major differentiator and one of the most prized assets for a company going forward. And, I’m not talking about launch videos or pouring money primarily for “aes…
MindTopo reveals VLMs’ spatial reasoning abilities
A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial re…
Introducing Mistral Small 4 | Mistral AI
The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.
AMIE, our research medical AI system, demonstrates real-time clinical video consultation capabilities in a first-of-its-kind study.
Google introduces AMIE for real-time clinical video consultations in simulated settings.
Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Toward…
Pomelli is our Google Labs experiment popular with a growing number of small businesses. Why? I…
Pomelli is our Google Labs experiment popular with a growing number of small businesses. Why? It's one of the easiest ways to make your products look good in any setting. With this launch, turn those amazing photoshoots…
Added a short instruction to our shared AGENTS MD file to upload videos to each PR that changes…
Added a short instruction to our shared AGENTS MD file to upload videos to each PR that changes UI state. https://t.co/BjpJL1qFSR https://t.co/09wZoz7X0J
Grok Bot is a fantastic release.. the UX, the design, the onboarding :chefs-kiss: But I go back…
Grok Bot is a fantastic release.. the UX, the design, the onboarding :chefs-kiss: But I go back and forth on this a lot: Will users (personal and business) want one super agent that holds all the context or sub agents t…
[AINews] Zawinski's Law of MultiAgents
a quiet day lets us find some connections among recent themes
story time.. in 2023 (my previous role) I remember the first time some customers volunteered th…
story time.. in 2023 (my previous role) I remember the first time some customers volunteered their prompt logs to us. We were surprised by the number of “build me an app for X” prompts. felt wildly ambitious at the time…
A great way to learn design: Give Codex a well-designed website, ask it to analyze what makes t…
A great way to learn design: Give Codex a well-designed website, ask it to analyze what makes the design great, then ask it to take a complete screenshot of the website & add annotations on the image that breaks down wh…
[AINews] Jeff, Sanjay, Oriol, and Quoc depart DeepMind; Demis to Chair; Koray to SVP — what is going on at GDM???
The end of an era.
One of my favorite videos lately! So many practical design tips; my favorite is “Just reduce th…
One of my favorite videos lately! So many practical design tips; my favorite is “Just reduce the font weights” and it magically makes the design look better https://t.co/4XsIiRIOV4
Exactly this. The OP generated some blurry images of people dressed in business casual having d…
Exactly this. The OP generated some blurry images of people dressed in business casual having dinner at an expensive restaurant. This is not cool. It has never been cool. https://t.co/QkX4tJnBPz
Grok Imagine Image 2.0 on Vercel AI Gateway Excellent 🖼️ model, #2 already on https://t.co/OSJ…
Grok Imagine Image 2.0 on Vercel AI Gateway Excellent 🖼️ model, #2 already on https://t.co/OSJc7zJ3gi https://t.co/MdjLBUeyxj
This paper gives a fancy name to a problem you can already feel in your bones "The Tragedy of t…
This paper gives a fancy name to a problem you can already feel in your bones "The Tragedy of the Cognitive Commons." Checking AI output requires deep expertise. Deep expertise comes from doing grunt work for years. And…
/human-review now has 500+ GitHub stars! I used it all day yesterday to edit some HTML and made…
/human-review now has 500+ GitHub stars! I used it all day yesterday to edit some HTML and made a few improvements. Now you can: 1. Make bulleted and numbered lists by typing “-” or “1.” 2. Add links by selecting text a…
How to build long-horizon AI agents: behavior specs, ontologies, process supervision - my conve…
How to build long-horizon AI agents: behavior specs, ontologies, process supervision - my conversation with @mitch_troy, co-founder of @trybasis 01:09 Why Everyone at Basis Was Whispering to AI when @steph_palazzolo wal…
How to use my new /human-review skill to edit HTML and Markdown files directly: 1. Tell Codex o…
How to use my new /human-review skill to edit HTML and Markdown files directly: 1. Tell Codex or Claude Code: “Install this: https://t.co/VDArIG0lxf” 2. Type “/human-review (your doc name)” 3. It’ll open a visual editor…
I gave codex a video-enabled remote KVM so it can automate e2e test the iMessage-integration on…
I gave codex a video-enabled remote KVM so it can automate e2e test the iMessage-integration on OpenClaw. (iMessage is unreliable in VMs, and certain features such as read receipts require SIP to be disabled)
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Baseten just raised a $13B Series F and is now one of the leading kings of inference engineering. We go into everything you need to know for autoregressive and diffusion engineering.
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
[AINews] not much happened today
apart from DeepSeek V4-Flash 0731, a quiet day.
The Biggest AI Deployment Nobody Talks About | Samsara CEO Sanjit Biswas
Speaker 1 | 00:00 - 00:21 These are not the tokens you're gonna find online. Like, you can't crawl Reddit and find out about what happened on a construction site. The Samsara system, in a given day, we're driving 99% of The US roads, usual…
Some observations from our Demo Faire yesterday: 1/ There is a tremendous amount of interest fr…
Some observations from our Demo Faire yesterday: 1/ There is a tremendous amount of interest from the larger investor community for frontier technology. Robots. Drones. Semiconductors. Bring it all. 2/ The interest in F…
[AINews] Much ado about Open Weights
Everyone is writing a lot, but only Kimi K3 shipped today
I also built Tastemaker to see if I could turn a rough idea into a useful product with Claude D…
I also built Tastemaker to see if I could turn a rough idea into a useful product with Claude Design and Claude Code. I made a video walking through my full process, including: 1. Creating a design.md and HTML spec 2. P…
I got tired of using IMDb to rate movies and TV shows, and Letterboxd doesn’t support video gam…
I got tired of using IMDb to rate movies and TV shows, and Letterboxd doesn’t support video games. So I built Tastemaker to let anyone curate their favorite movies, TV shows, and games all in one beautiful profile. You…
Great piece and vision for AI from Zuckerberg https://t.co/kO0oQbx3CC https://t.co/XjwKwvBZIu
Great piece and vision for AI from Zuckerberg https://t.co/kO0oQbx3CC https://t.co/XjwKwvBZIu
[AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable)
ain't nobody beats Anthropic at distilling Fable!
I asked @jxnlco (DevEx at OpenAI): “What’s the craziest thing Codex did for you?” Here’s his an…
I asked @jxnlco (DevEx at OpenAI): “What’s the craziest thing Codex did for you?” Here’s his answer: “I was on a bike ride when a coworker asked me to fix a launch video. So I connected remotely from my phone and asked…
[AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model
A HUGE win for BFL!
Inside the Model Factory — Eiso Kant, Poolside AI
Poolside's co-CEO on how his small team of top researchers built a model factory capable of training Laguna S - a 118B MOE beating Thinky's ~1T open weights model... and this is just the beginning.
One of our engineers built an interactive math art generator using the new 3.6 Flash model. You…
One of our engineers built an interactive math art generator using the new 3.6 Flash model. You can customize speed, colors, and geometry parameters on the fly, then export the design directly to a 3D-printable STL file…
Create, edit and star in videos with two Google Vids updates
Gemini Omni and personal avatars in Google Vids make video creation easier than ever.
[AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing
a great week for open models continues.
Use one agent to do the work and another to review it against a rubric. @trq212 explains how th…
Use one agent to do the work and another to review it against a rubric. @trq212 explains how this works for video shorts: “For something like, ‘Is this a good video short or not?’ you don’t have a deterministic answer.…
The big lesson from AI is that everything is code. A slide deck is code. Design is code. That c…
The big lesson from AI is that everything is code. A slide deck is code. Design is code. That cool promo video? Code. Excel automation? Code. The universe? Probably made of code too.
🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences
Lila is betting that science, not the internet, is the last untapped source of training data. We went to find out what that actually looks like in a room full of robots.
Whenever someone asks me "I'm just getting started with posting content on social media; what s…
Whenever someone asks me "I'm just getting started with posting content on social media; what should I talk about?" My answer is: What are the top 3 questions that you get asked most frequently by friends/acquaintances/…
Celebrating 25 years of visual search innovation
Google Images is turning 25. Here’s a look back at some major milestones — and new ways to explore and create visual content.