Interestana
Home/Topics/Machine Learning
๐Ÿค–Topic

Machine Learning

2 articles curated by AI agents. Last updated Just now.

Machine learning is seeing advancements in multimodal embedding models, with Google DeepMind releasing EmbeddingGemma 2. Research is also pushing the boundaries of robotics learning pipelines and the development of domain-agnostic world models. Furthermore, there's ongoing discussion about the reasoning capabilities of current Large Language Models (LLMs).

Machine Learning: Questions & Answers

Answers synthesised from 12 recent sources ยท updated 4h ago

What are the latest developments in multimodal embedding models?

Google DeepMind has released EmbeddingGemma 2, an open-source multimodal embedding model. Built on the Gemma 4 architecture, it features 740 million parameters and can convert text, code, images, video, and audio into a unified 768-dimensional vector space.

What is JEPA-Anything and what does it do?

JEPA-Anything is a novel domain-agnostic framework for constructing world models, introduced by researchers from PhAI Labs, CUHK, Fudan, Stanford, Oxford, and Princeton. This approach moves beyond domain-specific world models by using a single recipe for seven different fields.

What are the recent achievements of Google DeepMind's Nemotron model family?

Google DeepMind's Nemotron model family has achieved gold medal-level performance in two prestigious international competitions: the International Olympiad in Informatics (IOI) and the International Mathematical Olympiad (IMO).

What is Laya and what are its key features?

Laya is an open-source decision engine released by Convai Innovations, designed for rapid, calibrated probability generation. It utilizes a 421-million-parameter encoder that processes text and questions in a single forward pass.

What is the current consensus on the reasoning abilities of Large Language Models (LLMs)?

Thore Graepel, a former core member of the AlphaGo team at Google DeepMind, states that current Large Language Models (LLMs) do not possess true reasoning abilities. This is a critical distinction from earlier AI systems like AlphaGo.

How is Google Research advancing Federated Learning?

Google Research has developed a next-generation Federated Learning (FL) system that uses Trusted Execution Environments (TEEs) to provide externally verifiable central differential privacy (DP) guarantees. This system addresses a trust gap in previous FL implementations.

MIT Technology Review2h ago3 min read
The Download: AI roadblocks for humanoids and portable rubber dams

The current hype surrounding humanoid robots, fueled by the success of AI models like ChatGPT and Claude, is met with skepticism from many researchers. These experts argue that the intelligence developed for language and image processing may not be sufficient to navigate the "infinite variability of the physical world." A key concern is the conflation of general-purpose humanoids with generalist machines, suggesting that current AI paradigms might not be adequate for creating truly versatile robots. This has led to a debate within robotics labs about whether existing AI advancements are enough or if a fundamentally new approach is necessary to perfect all-purpose humanoids. The development of truly general-purpose robots is a complex undertaking that extends beyond current AI capabilities in language and image recognition, requiring a deeper understanding of physical interaction and adaptation. In parallel, the climate technology sector is seeing innovation in flood prevention. WaveSave has been recognized as one of MIT Technology Review's 10 Climate Tech Companies to Watch in 2026. The company has developed a portable rubber dam, named the SlamDam, designed for rapid deployment during flood events. This ingeniously simple tool can be unfurled along riverbanks or shorelines and then inflated with water, creating a barrier up to 1.3 meters high. The SlamDam aims to protect nearby properties and structures from rising water levels. WaveSave intends to expand the application and reach of its innovative flood protection solution, addressing the increasing risks associated with sea-level rise and worsening flood conditions. This development highlights a practical application of technology to mitigate the impacts of climate change, offering a tangible solution for communities vulnerable to flooding. The broader implications of these developments point to distinct challenges and opportunities in technological advancement. While AI's potential in abstract domains like language and image generation is rapidly progressing, its translation to complex physical systems like humanoid robots faces significant hurdles. This suggests a potential divergence in AI development trajectories, with some areas seeing exponential growth while others require more foundational research and a different set of technological breakthroughs. Conversely, the success of companies like WaveSave demonstrates that practical, albeit less complex, technological solutions can emerge to address pressing environmental challenges. The focus on tangible, deployable technologies for climate resilience contrasts with the more speculative, long-term ambitions in advanced robotics. The distinction between AI's current strengths and the requirements for physical world mastery remains a critical area of research and development, influencing the timeline and nature of future robotic applications. The development of portable flood defenses also underscores the growing need for adaptive infrastructure in the face of a changing climate, showcasing how innovation can provide immediate benefits.

MarkTechPost5h ago3 min read
NVIDIA PivotOPD Teaches Multi-Turn AI Agents to Recover From Pivotal Mistakes

NVIDIA researchers, in collaboration with Princeton University and the University of Maryland, have introduced PivotOPD, a novel on-policy distillation method designed to train multi-turn large language model (LLM) agents. This technique specifically focuses on teaching an agent to avoid its most detrimental early mistakes and, crucially, to recover effectively when such errors inevitably occur. PivotOPD demonstrated superior performance against 13 established baselines across three benchmark environments: ALFWorld, WebShop, and Search-based QA, when applied to Qwen3-1.7B and Qwen3-8B student models. The core finding is that the ability to recover from errors is a learnable trait, and traditional on-policy distillation (OPD) methods typically fail to impart this critical capability. PivotOPD is a training methodology rather than a new model itself. It has been tested on student models including Qwen3-1.7B, Qwen3-8B, and a Nemotron-3.5-SFT student model evaluated on the SWE-Bench Verified dataset. The training process was conducted on NVIDIA H100 nodes. A significant advantage of PivotOPD is that it introduces no additional inference cost, meaning that agents trained with this method can be deployed and run on any hardware compatible with their base models. In performance evaluations, PivotOPD achieved the top average results across all eight per-benchmark metrics against 13 baselines, with results averaged over three random seeds. Specifically, PivotOPD-trained agents were able to recover from 72.7% of replayed pivotal mistakes, a substantial improvement over the 20.3% recovery rate achieved by standard OPD. The lowest recovery rate observed was 55.9% on ALFWorld's "Look" tasks using the 1.7B student model, compared to 83.9% for a baseline method referred to as SOD. A pivotal mistake in the context of multi-turn agents is defined as an action that either prolongs the shortest possible path to task completion or renders the task unsolvable. ALFWorld's symbolic oracle quantifies this at each interaction turn. Analysis of Qwen3-8B, Qwen3-30B-A3B, and Qwen3-235B-A22B models revealed that 59% of failed task executions (155 out of 262 instances) contained at least one pivotal mistake. These critical errors often occurred early in the task, with a median arrival turn between 8 and 12 out of a total of 30 turns. Following such a mistake, agents frequently wasted an additional 18 to 21 turns without successfully recovering. When pivotal mistakes in Qwen3-8B failures were corrected during replays, the success rate increased from a mere 8% to 59%. Even when the mistake was left in place but the agent was forced to take the correct action for the subsequent two turns, the success rate reached 58%, underscoring the impact of correcting or mitigating early errors. The research also investigates why standard on-policy distillation methods fall short in teaching recovery. Standard OPD managed to reduce the held-out failure rate from 79% to 56%. However, failures that occurred specifically after a pivotal turn remained a significant issue, indicating that this method does not adequately address the consequences of such critical errors. The effectiveness of PivotOPD is contingent on the availability of replayable environments and a teacher model whose critical action choices, or "pivots," align with the oracle's decisions in at least 77.8% of failed rollouts. This highlights the importance of a high-quality teacher signal for successful on-policy distillation focused on error recovery.