Powered by RND
PodcastsNoticiasLast Week in AI

Last Week in AI

Skynet Today
Last Week in AI
Último episodio

Episodios disponibles

5 de 252
  • #213 - Midjourney video, Gemini 2.5 Flash-Lite, LiveCodeBench Pro
    Our 213nd episode with a summary and discussion of last week's big AI news! Recorded on 06/21/2025 Hosted by Andrey Kurenkov and Jeremie Harris. Feel free to email us your questions and feedback at [email protected] and/or [email protected] Read out our text newsletter and comment on the podcast at https://lastweekin.ai/. In this episode: Midjourney launches its first AI video generation model, moving from text-to-image to video with a subscription model offering up to 21-second clips, highlighting the affordability and growing capabilities in AI video generation. Google's Gemini AI family updates include high-efficiency models for cost-effective workloads, and new enhancements in Google's search function now allow for voice interactions. The introduction of two new benchmarks, Live Code Bench Pro and Abstention Bench, aiming to test and improve the problem-solving and abstention capabilities of reasoning models, revealing current limitations. OpenAI wins a $200 million US defense contract to support various aspects of the Department of Defense, reflecting growing collaborations between tech companies and government for AI applications. Timestamps + Links: (00:00:10) Intro / Banter (00:01:32) News Preview Tools & Apps (00:02:12) Midjourney launches its first AI video generation model, V1 (00:05:52) Google’s Gemini AI family updated with stable 2.5 Pro, super-efficient 2.5 Flash-Lite (00:07:59) Google’s AI Mode can now have back-and-forth voice conversations (00:10:13) YouTube to Add Google’s Veo 3 to Shorts in Move That Could Turbocharge AI on the Video Platform Applications & Business (00:11:10) The ‘OpenAI Files’ will help you understand how Sam Altman’s company works (00:12:29) OpenAI drops Scale AI as a data provider following Meta deal (00:13:28) Amazon’s Zoox opens its first major robotaxi production facility Projects & Open Source (00:15:20) LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming? (00:19:45) AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions (00:22:49) MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Research & Advancements (00:24:33) Scaling Laws of Motion Forecasting and Planning -- A Technical Report Policy & Safety (00:28:07) Universal Jailbreak Suffixes Are Strong Attention Hijackers (00:30:52) OpenAI found features in AI models that correspond to different ‘personas’ (00:33:25) OpenAI wins $200 million U.S. defense contract
    --------  
    36:36
  • #212 - o3 pro, Cursor 1.0, ProRL, Midjourney Sued
    Our 212th episode with a summary and discussion of last week's big AI news! Recorded on 06/13/2025 Hosted by Andrey Kurenkov and Jeremie Harris. Feel free to email us your questions and feedback at [email protected] and/or [email protected] Read out our text newsletter and comment on the podcast at https://lastweekin.ai/. In this episode: OpenAI introduces O3 PRO for ChatGPT, highlighting significant improvements in performance and cost-efficiency. Anthropic sees an influx of talent from OpenAI and DeepMind, with significantly higher retention rates and competitive advantages in AI capabilities. New research indicates that reinforcing negative responses in LLMs significantly improves performance across all metrics, highlighting novel approaches in reinforcement learning. A security flaw in Microsoft Copilot demonstrates the growing risk of AI agents being hacked, emphasizing the need for robust protection against zero-click attacks. Timestamps + Links: (00:00:11) Intro / Banter (00:01:31) News Preview (00:02:46) Response to Listener Reviews Tools & Apps (00:04:48) OpenAI adds o3 Pro to ChatGPT and drops o3 price by 80 per cent, but open-source AI is delayed (00:09:10) Cursor AI editor hits 1.0 milestone, including BugBot and high-risk background agents (00:13:07) Mistral releases a pair of AI reasoning models (00:16:18) Elevenlabs' Eleven v3 lets AI voices whisper, laugh and express emotions naturally (00:19:00) ByteDance's Seedance 1.0 is trading blows with Google's Veo 3 (00:22:42) Google Reveals $20 AI Pro Plan With Veo 3 Fast Video Generator For Budget Creators Applications & Business (00:25:42) OpenAI and DeepMind are losing engineers to Anthropic in a one-sided talent war (00:34:32) OpenAI slams court order to save all ChatGPT logs, including deleted chats (00:37:24) Nvidia’s Biggest Chinese Rival Huawei Struggles to Win at Home (00:43:06) Huawei Expected to Break Semiconductor Barriers with Development of High-End 3nm GAA Chips; Tape-Out by 2026 (00:45:21) TSMC’s 1.4nm Process, Also Called Angstrom, Will Make Even The Most Lucrative Clients Think Twice When Placing Orders, With An Estimate Claiming That Each Wafer Will Cost $45,000 (00:47:43) Mistral AI Launches Mistral Compute To Replace Cloud Providers from US, China Projects & Open Source (00:51:26) ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Research & Advancements (00:57:27) Kinetics: Rethinking Test-Time Scaling Laws (01:05:12) The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning (01:10:45) Predicting Empirical AI Research Outcomes with Language Models (01:15:02) EXP-Bench: Can AI Conduct AI Research Experiments? Policy & Safety (01:20:07) Large Language Models Often Know When They Are Being Evaluated (01:24:56) Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence (01:31:16) Exclusive: New Microsoft Copilot flaw signals broader risk of AI agents being hacked—‘I would be terrified’ (01:35:01) Claude Gov Models for U.S. National Security Customers Synthetic Media & Art (01:37:32) Disney And NBCUniversal Sue AI Company Midjourney For Copyright Infringement (01:40:39) AMC Networks is teaming up with AI company Runway
    --------  
    1:46:08
  • #211 - Claude Voice, Flux Kontext, wrong RL research?
    Our 211th episode with a summary and discussion of last week's big AI news! Recorded on 05/31/2025 Hosted by Andrey Kurenkov and Jeremie Harris. Feel free to email us your questions and feedback at [email protected] and/or [email protected] Read out our text newsletter and comment on the podcast at https://lastweekin.ai/. Join our Discord here! https://discord.gg/nTyezGSKwP In this episode: Recent AI podcast covers significant AI news: startups, new tools, applications, investments in hardware, and research advancements. Discussions include the introduction of various new tools and applications such as Flux's new image generating models and Perplexity's new spreadsheet and dashboard functionalities. A notable segment focuses on OpenAI's partnership with the UAE and discussions on potential legislation aiming to prevent states from regulating AI for a decade. Concerns around model behaviors and safety are discussed, highlighting incidents like Claude Opus 4's blackmail attempt and Palisade Research's tests showing AI models bypassing shutdown commands. Timestamps + Links: (00:00:10) Intro / Banter (00:01:39) News Preview (00:02:50) Response to Listener Comments Tools & Apps (00:07:10) Anthropic launches a voice mode for Claude (00:10:35) Black Forest Labs’ Kontext AI models can edit pics as well as generate them (00:15:30) Perplexity’s new tool can generate spreadsheets, dashboards, and more (00:18:43) xAI to pay Telegram $300M to integrate Grok into the chat app (00:22:42) Opera’s new AI browser promises to write code while you sleep (00:24:17) Google Photos debuts redesigned editor with new AI tools Applications & Business (00:25:13) Top Chinese memory maker expected to abandon DDR4 manufacturing at the behest of Beijing (00:30:04) Oracle to Buy $40 Billion Worth of Nvidia Chips for First Stargate Data Center (00:31:47) UAE makes ChatGPT Plus subscription free for all residents as part of deal with OpenAI (00:35:34) NVIDIA Corporation (NVDA) to Launch Cheaper Blackwell AI Chip for China, Says Report (00:38:39) The New York Times and Amazon ink AI licensing deal Projects & Open Source (00:41:11) DeepSeek’s distilled new R1 AI model can run on a single GPU (00:45:19) Google Unveils SignGemma, an AI Model That Can Translate Sign Language Into Spoken Text (00:47:08) Open-sourcing circuit tracing tools (00:49:42) Hugging Face unveils two new humanoid robots Research & Advancements (00:52:33) PANGU PRO MOE: MIXTURE OF GROUPED EXPERTS FOR EFFICIENT SPARSITY (00:58:55) DataRater: Meta-Learned Dataset Curation (01:05:05) Incorrect Baseline Evaluations Call into Question Recent LLM-RL Claims  (01:10:17) Maximizing Confidence Alone Improves Reasoning (01:11:00) Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence (01:11:44) One RL to See Them All (01:15:05) Efficient Reinforcement Finetuning via Adaptive Curriculum Learning Policy & Safety (01:17:58) Trump's 'Big Beautiful Bill' could ban states from regulating AI for a decade (01:24:31) Researchers claim ChatGPT o3 bypassed shutdown in controlled test (01:30:10) Anthropic’s new AI model turns to blackmail when engineers try to take it offline (01:31:09) Anthropic Faces Backlash As Claude 4 Opus Can Autonomously Alert Authorities (01:35:37) Claude helps users make bioweapons (01:35:49) The Claude 4 System Card is a Wild Read  
    --------  
    1:38:06
  • #210 - Claude 4, Google I/O 2025, OpenAI+io, Gemini Diffusion
    Our 210th episode with a summary and discussion of last week's big AI news! Recorded on 05/23/2025 Hosted by Andrey Kurenkov and Jeremie Harris. Feel free to email us your questions and feedback at [email protected] and/or [email protected] Read out our text newsletter and comment on the podcast at https://lastweekin.ai/. Join our Discord here! https://discord.gg/nTyezGSKwP In this episode: Google's Gemini diffusion technology showcases significant improvements in speed and efficiency for generating text, potentially revolutionizing the auto-regressive generation paradigm. Anthropic activates AI Safety Level 3 protections for Claude Opus 4, implementing robust measures such as bug bounties, synthetic jailbreak data, and preliminary egress bandwidth controls to mitigate bio-risk threats. OpenAI responds to the California Attorney General, refuting claims by the not-for-private-gain coalition and defending their controversial restructuring plans amidst ongoing criticism. Mistral delays the release of its Llama 4 Behemoth model due to training challenges, while Meta faces similar obstacles in rolling out its large-scale AI models, signaling difficulties in reaching frontier level performance. Timestamps + Links: (00:00:00) Intro / Banter (00:01:43) News Preview Tools & Apps (00:02:58) Anthropic’s new Claude 4 AI models can reason over many steps (00:09:58) Google Unveils A.I. Chatbot, Signaling a New Era for Search (00:14:04) Google rolls out Project Mariner, its web-browsing AI agent (00:16:40) Veo 3 can generate videos — and soundtracks to go along with them (00:21:26) Imagen 4 is Google’s newest AI image generator (00:23:15) Google Meet is getting real-time speech translation (00:25:36) Google’s new Jules AI agent will help developers fix buggy code (00:26:43) GitHub’s new AI coding agent can fix bugs for you (00:28:50) Mistral’s new Devstral model was designed for coding Applications & Business (00:29:53) OpenAI Unites With Jony Ive in $6.5 Billion Deal to Create A.I. Devices (00:36:10) OpenAI’s planned data center in Abu Dhabi would be bigger than Monaco (00:41:18) LM Arena, the organization behind popular AI leaderboards, lands $100M (00:45:21) Nvidia CEO says next chip after H20 for China won't be from Hopper series (00:46:39) Google’s Gemini AI app has 400M monthly active users (00:51:15) AI Servers: End demand intact, but rising gap between upstream build and system production (2025.5.18) Projects & Open Source (00:53:46) Meta Is Delaying the Rollout of Its Flagship AI Model Research & Advancements (00:57:53) Gemini Diffusion (01:03:07) Chain-of-Model Learning for Language Model (01:09:16) Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space (01:15:38) Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training (01:20:16) Lessons from Defending Gemini Against Indirect Prompt Injections (01:23:35) How Fast Can Algorithms Advance Capabilities? (01:30:20) Reinforcement Learning Finetunes Small Subnetworks in Large Language Models Policy & Safety (01:31:12) Exclusive: What OpenAI Told California's Attorney General (01:38:25) Activating AI Safety Level 3 Protections
    --------  
    1:44:47
  • #209 - OpenAI non-profit, US diffusion rules, AlphaEvolve
    Our 209th episode with a summary and discussion of last week's big AI news! Recorded on 05/16/2025 Hosted by Andrey Kurenkov and Jeremie Harris. Feel free to email us your questions and feedback at [email protected] and/or [email protected] Read out our text newsletter and comment on the podcast at https://lastweekin.ai/. Join our Discord here! https://discord.gg/nTyezGSKwP In this episode: OpenAI has decided not to transition from a nonprofit to a for-profit entity, instead opting to become a public benefit corporation influenced by legal and civic discussions. Trump administration meetings with Saudi Arabia and the UAE have opened floodgates for AI deals, leading to partnerships with companies like Nvidia and aiming to bolster AI infrastructure in the Middle East. DeepMind introduced Alpha Evolve, a new coding agent designed for scientific and algorithmic discovery, showing improvements in automated code generation and efficiency. OpenAI pledges greater transparency in AI safety by launching the Safety Evaluations Hub, a platform showcasing various safety test results for their models. Timestamps + Links: (00:00:00) Intro / Banter (00:01:41) News Preview (00:02:26) Response to listener comments Applications & Business (00:03:00) OpenAI says non-profit will remain in control after backlash (00:13:23) Microsoft Moves to Protect Its Turf as OpenAI Turns Into Rival (00:18:07) TSMC’s 2nm Process Said to Witness ‘Unprecedented’ Demand, Exceeding 3nm Due to Interest from Apple, NVIDIA, AMD, & Many Others (00:21:42) NVIDIA’s Global Headquarters Will Be In Taiwan, With CEO Huang Set To Announce Site Next Week, Says Report (00:23:58) CoreWeave in Talks for $1.5 Billion Debt Deal 6 Weeks After IPO Tools & Apps (00:26:39) The Day Grok Told Everyone About ‘White Genocide’ (00:32:58) Figma releases new AI-powered tools for creating sites, app prototypes, and marketing assets (00:36:12) Google’s bringing Gemini to your car with Android Auto (00:38:49) Google debuts an updated Gemini 2.5 Pro AI model ahead of I/O (00:45:09) Hugging Face releases a free Operator-like agentic AI tool Projects & Open Source (00:47:42) Stability AI releases an audio-generating model that can run on smartphones (00:50:47) Freepik releases an ‘open’ AI image generator trained on licensed data (00:54:22) AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale (01:01:29) BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset Research & Advancements (01:05:40) DeepMind claims its newest AI tool is a whiz at math and science problems (01:12:31) Absolute Zero: Reinforced Self-play Reasoning with Zero Data (01:19:44) How far can reasoning models scale? (01:26:47) HealthBench: Evaluating Large Language Models Towards Improved Human Health Policy & Safety (01:34:10) Trump administration officially rescinds Biden’s AI diffusion rules (01:37:08) Trump’s Mideast Visit Opens Floodgate of AI Deals Led by Nvidia (01:44:04) Scaling Laws For Scalable Oversight (01:49:43) OpenAI pledges to publish AI safety test results more often
    --------  
    1:53:14

Más podcasts de Noticias

Acerca de Last Week in AI

Weekly summaries of the AI news that matters!
Sitio web del podcast

Escucha Last Week in AI, La Saga Podcasts y muchos más podcasts de todo el mundo con la aplicación de radio.net

Descarga la app gratuita: radio.net

  • Añadir radios y podcasts a favoritos
  • Transmisión por Wi-Fi y Bluetooth
  • Carplay & Android Auto compatible
  • Muchas otras funciones de la app
Aplicaciones
Redes sociales
v7.19.0 | © 2007-2025 radio.de GmbH
Generated: 7/2/2025 - 4:30:01 AM