Interestana
Home/Topics/Generative AI
🤖Topic

Generative AI

3 articles curated by AI agents. Last updated Just now.

Generative AI is seeing rapid advancements with new model releases focused on efficiency, multimodal capabilities, and specialized tasks. Companies like Perplexity AI, Anthropic, and Liquid AI are introducing models with varying parameter sizes, context windows, and pricing structures to cater to diverse application needs.

Generative AI: Questions & Answers

Answers synthesised from 3 recent sources · updated 5h ago

What new multimodal embedding models has Perplexity AI released?

Perplexity AI has released pplx-embed-v2-late, a suite of ColBERT-style multimodal embedding models. This release includes a 0.6 billion parameter model optimized for speed and cost-efficiency, and a 9 billion parameter model designed for maximum performance.

What are the key features of Anthropic's Claude Haiku 5.5?

Anthropic released Claude Haiku 5.5 this week, positioning it as its most affordable and rapid small-scale AI model. It is engineered for high-volume tasks and features a 1 million token context window, priced at $0.10 per million input tokens.

What distinguishes Liquid AI's d1 decision models?

Liquid AI has released open-weight multimodal models, d1-3B and d1-omni-600M, which are engineered to provide calibrated, typed answers in a single forward pass. These models produce zero output tokens, differentiating them from traditional generative models.

What are the different sizes of Perplexity AI's new embedding models?

Perplexity AI's new suite of ColBERT-style multimodal embedding models, pplx-embed-v2-late, offers two sizes: a 0.6 billion parameter model and a 9 billion parameter model.

What types of tasks is Anthropic's Claude Haiku 5.5 designed for?

Claude Haiku 5.5 is engineered for high-volume tasks such as text summarization, data compaction, classification, and the operation of sub-agents within larger AI systems.

What are the names of Liquid AI's newly released open-weight multimodal models?

Liquid AI has released two open-weight multimodal models: d1-3B and d1-omni-600M, as part of its d1 decision model family.

MIT Technology Review4h ago4 min read
AI breakthroughs in robotics won’t change your life any time soon

Recent advancements in artificial intelligence have fueled significant optimism regarding the future of robotics, particularly humanoid robots. Elon Musk, CEO of Tesla, has been a vocal proponent, envisioning his company's Optimus robot as a transformative product capable of automating vast amounts of human labor. Musk stated in July that Optimus robots will eventually possess "human and then superhuman dexterity," and he predicts they could be available to the public by the end of 2027. He further projects that these robots could automate tasks ranging from hauling sheet metal to folding laundry, with a potential cost as low as $20,000 each. Musk believes Optimus will become "not just Tesla’s biggest product ever, but probably the biggest product ever," initially deployed on factory floors and later in homes. This enthusiasm is echoed by other prominent figures in the tech industry. Marc Andreessen, cofounder and general partner at venture capital firm Andreessen Horowitz, has suggested that robotics could emerge as the "biggest industry in the history of the planet." Jensen Huang, CEO of Nvidia, stated in January that humanoid robots are expected to achieve human-level abilities within the current year. Financial projections also reflect this optimism, with Morgan Stanley forecasting that the number of robots resembling and acting like humans could approach 1 billion by 2050, potentially creating a market valued at over $5 trillion. Currently, Tesla's Optimus is most frequently observed performing tasks such as distributing food and beverages at company events, demonstrating its nascent capabilities. However, the path from these impressive demonstrations to widespread integration into daily life is fraught with challenges. The current iterations of humanoid robots, while showcasing progress, still exhibit limitations. Videos have depicted robots struggling with tasks like ironing shirts or maintaining balance when handing out items, such as falling backward while distributing water bottles. These instances highlight the significant gap between current performance and the sophisticated, reliable dexterity required for domestic or complex industrial applications. The development of robust AI that can navigate unpredictable real-world environments, adapt to novel situations, and perform delicate manipulations with human-like precision remains a formidable technical hurdle. Furthermore, the economic and societal implications of such widespread automation are complex and require careful consideration. While the potential cost reduction to $20,000 per robot is cited as a driver for mass adoption, the initial investment in research, development, manufacturing, and deployment will be substantial. The integration of a billion humanoid robots into society by 2050, as projected by Morgan Stanley, would necessitate significant infrastructure changes, regulatory frameworks, and societal adjustments to address potential job displacement and ethical concerns. The collaborative efforts between institutions like MIT Technology Review and research foundations such as Aventine aim to explore these evolving dynamics, underscoring that while the technological trajectory is promising, the timeline for transformative impact on everyday life is likely longer than some optimistic predictions suggest.

MarkTechPost5h ago3 min read
NVIDIA PivotOPD Teaches Multi-Turn AI Agents to Recover From Pivotal Mistakes

NVIDIA researchers, in collaboration with Princeton University and the University of Maryland, have introduced PivotOPD, a novel on-policy distillation method designed to train multi-turn large language model (LLM) agents. This technique specifically focuses on teaching an agent to avoid its most detrimental early mistakes and, crucially, to recover effectively when such errors inevitably occur. PivotOPD demonstrated superior performance against 13 established baselines across three benchmark environments: ALFWorld, WebShop, and Search-based QA, when applied to Qwen3-1.7B and Qwen3-8B student models. The core finding is that the ability to recover from errors is a learnable trait, and traditional on-policy distillation (OPD) methods typically fail to impart this critical capability. PivotOPD is a training methodology rather than a new model itself. It has been tested on student models including Qwen3-1.7B, Qwen3-8B, and a Nemotron-3.5-SFT student model evaluated on the SWE-Bench Verified dataset. The training process was conducted on NVIDIA H100 nodes. A significant advantage of PivotOPD is that it introduces no additional inference cost, meaning that agents trained with this method can be deployed and run on any hardware compatible with their base models. In performance evaluations, PivotOPD achieved the top average results across all eight per-benchmark metrics against 13 baselines, with results averaged over three random seeds. Specifically, PivotOPD-trained agents were able to recover from 72.7% of replayed pivotal mistakes, a substantial improvement over the 20.3% recovery rate achieved by standard OPD. The lowest recovery rate observed was 55.9% on ALFWorld's "Look" tasks using the 1.7B student model, compared to 83.9% for a baseline method referred to as SOD. A pivotal mistake in the context of multi-turn agents is defined as an action that either prolongs the shortest possible path to task completion or renders the task unsolvable. ALFWorld's symbolic oracle quantifies this at each interaction turn. Analysis of Qwen3-8B, Qwen3-30B-A3B, and Qwen3-235B-A22B models revealed that 59% of failed task executions (155 out of 262 instances) contained at least one pivotal mistake. These critical errors often occurred early in the task, with a median arrival turn between 8 and 12 out of a total of 30 turns. Following such a mistake, agents frequently wasted an additional 18 to 21 turns without successfully recovering. When pivotal mistakes in Qwen3-8B failures were corrected during replays, the success rate increased from a mere 8% to 59%. Even when the mistake was left in place but the agent was forced to take the correct action for the subsequent two turns, the success rate reached 58%, underscoring the impact of correcting or mitigating early errors. The research also investigates why standard on-policy distillation methods fall short in teaching recovery. Standard OPD managed to reduce the held-out failure rate from 79% to 56%. However, failures that occurred specifically after a pivotal turn remained a significant issue, indicating that this method does not adequately address the consequences of such critical errors. The effectiveness of PivotOPD is contingent on the availability of replayable environments and a teacher model whose critical action choices, or "pivots," align with the oracle's decisions in at least 77.8% of failed rollouts. This highlights the importance of a high-quality teacher signal for successful on-policy distillation focused on error recovery.

MarkTechPost8h ago3 min read
Perplexity AI Releases pplx-embed-v2-late: A 0.6B Edge Model and a 9B Model Scoring 92.4% on MADQA

Perplexity AI released pplx-embed-v2-late, a new suite of ColBERT-style multimodal embedding models, on an unspecified date, offering two distinct sizes: a 0.6 billion parameter model optimized for speed and cost-efficiency, and a 9 billion parameter model designed for maximum quality. Both models are capable of retrieving text, images, and rendered PDF pages, and importantly, they share a unified embedding space, allowing for cross-modal understanding. These models are available for self-hosting under the MIT license, which permits commercial use, and are accessible on Hugging Face. Perplexity has indicated plans for a hosted API endpoint, though it is not yet live. The 0.6B model is engineered to function as a lightweight query encoder, suitable for deployment on edge devices or laptops, utilizing approximately 340 million active parameters for image processing. This smaller model aims to maintain performance close to larger, 8 billion parameter rivals. A key feature highlighted is the ability to search a 9B index using 0.6B queries, which recovers about half of the quality gap observed in text retrieval at the reduced query cost. The models generate 128-dimensional token vectors, significantly narrower than competitors' vectors which range from 2,048 to 4,096 dimensions, representing a 16x to 32x reduction in dimensionality. However, the models present certain limitations. The architecture stores one vector per token, leading to index sizes that scale directly with document length. In terms of performance benchmarks, pplx-embed-v2-late is not the top performer on the ViDoRe v3 image retrieval task; Tencent's EVIE model reportedly scores higher. Furthermore, a single input cannot simultaneously contain both text and images. Perplexity has stated that all performance scores are self-reported, and the accompanying technical report has not yet been published. Technical specifications reveal that the pplx-embed-v2-late-0.6b model comprises 594 million total parameters, with approximately 240 million active parameters for text and 340 million for images. The base model is derived from Qwen3.5-0.8B, pruned to 12 text layers. The 9B variant has 9 billion total parameters (though Hugging Face lists it as 8B) and utilizes 7.4 billion parameters. Both output 128 dimensions per token. Estimated memory requirements for weights in bf16 precision are around 1.2 GB for the 0.6B model and 16 to 18 GB for the 9B model. The models require sentence-transformers version 6.0.0 or later and transformers version 5.4.0 or later. The published checkpoints are distributed in F32 format, which doubles their download size.