Open Source LLM

China approves RedNote’s AI translation along with 5 other models

In January, RedNote introduced a feature to automatically translate posts and comments. Photo: VCG via Getty Images

The internet regulator in Shanghai has approved six more generative artificial intelligence (GenAI) services for public release, including a translation tool from RedNote, a Chinese social platform that surged in popularity overseas last month amid uncertainties about TikTok ’s fate in the US.
The newly registered services also included an AI assistant for the upcoming Global Developer Conference, a three-day event in Shanghai that kicks off on Friday. Members of Chinese AI start-up

Open Source LLM

Linux Lazy Unmap Flush "LUF" Reducing TLB Shootdowns By 97%, Faster AI LLM Performance

SK has been working on a Linux kernel feature dubbed Lazy Unmap Flush "LUF" to defer TLB flushes until folios have been unmapped and freed are eventually allocated again.

This Lazy Unmap Flush work began after encountering a lot of migration overhead around TLB shootdowns on servers with tiered memory making use of CXL memory.

The end result is what is most interesting and important: the LUF patches yielded TLB shootdown interrupts being reduced by around 97%. Furthermore, the test program runtime of using Llama.cpp with a large language
Continue reading...

Open Source LLM

生成式 AI 推理應用新選擇 DeepSeek-R1 登陸 Amazon SageMaker JumpStart

LLM, AI Large Language Model concept. Businessman use tablet and laptop with LLM icons. A language model distinguished by its general-purpose language generation capability. Chat AI.

DeepSeek-R1 現已上架 Amazon SageMaker JumpStart ,用戶可以透過這些平台輕鬆部署並運行該模型,用於生成式 AI 應用的推理 (Inference)。無論是探索創新 AI 應用還是大規模部署解決


Open Source LLM

DeepSeek founder provides clue on start-up’s AI priorities in new technical study

A new DeepSeek study touts “native sparse attention” as a way to make artificial intelligence models more efficient when processing vast amounts of data. Photo: Shutterstock

DeepSeek has signalled its next development priorities in a new technical study, with founder and chief executive Liang Wenfeng among 15 co-authors, that delves on “native sparse attention” (NSA) – a system that is touted to make artificial intelligence (AI) models more efficient when processing vast amounts of data.
The study, titled “Native Sparse Attention:
Continue reading...

Open Source LLM

DeepSeek faster, cheaper: innovation speeds processing of long text 10 times, paper says

DeepSeek says its NSA method combines algorithm innovations with enhanced hardware to improve efficiency without sacrificing performance. Photo: AFP

Chinese AI start-up DeepSeek has unveiled a new technology that could allow next-generation language models to process very long text much faster and cheaper than traditional methods.
By training AI to focus on key information rather than every word, the company’s “native sparse attention” (NSA) method sped up long-text processing by up to 11 times
Continue reading...

Open Source LLM

How DeepSeek’s disruption could finally burst the US stock bubble

TV news at the Nasdaq headquarters in Times Square highlights the extent of the market drop on January 27 in New York, US. Photo: AFP

DeepSeek has unravelled the assumption that a few hyperscalers will dominate the artificial intelligence market, extracting payments from an AI-dependent global economy. The Chinese company’s open-source program lets businesses develop their own applications without having to pay anyone. The revenue assumption of hundreds of billions of dollars that underpins hyperscalers is looking like hot air.
The US economy depends on its buoyant stock market which, at twice the value of its gross domestic product

Open Source LLM

xAI Claims Grok 3 Is The “World’s Smartest AI,” Betting Markets Agree, But Experts Remain Split

After days of building up hype, xAI officially released its Grok 3 LLM on Monday in a live stream hosted by Elon Musk himself. While the AI company continues to tout the new LLM's capabilities as the best in-class, some experts are pointing out critical shortcomings in the released benchmarks. grok 3 is the world’s smartest AI now available to all Premium+ subscribers — Grok (@grok) February 18, 2025 To wit, as per xAI's post on X,

Open Source LLM

(PR) AMD & Nexa AI Reveal NexaQuant's Improvement of DeepSeek R1 Distill 4-bit Capabilities

Nexa AI, today, announced NexaQuants of two DeepSeek R1 Distills: The DeepSeek R1 Distill Qwen 1.5B and DeepSeek R1 Distill Llama 8B. Popular quantization methods like the llama.cpp based Q4 K M allow large language models to significantly reduce their memory footprint and typically offer low perplexity loss for dense models as a tradeoff. However, even low perplexity loss can result in a reasoning capability hit for (dense or MoE) models that use Chain of Thought traces. Nexa AI has stated

Open Source LLM

Modder crams LLM onto Pi Zero-powered USB stick, but it isn't fast enough to be practical

(Image credit: YouTube: Build with Binh)

Local LLM usage is on the rise, and with many setting up PCs or systems to run them, the idea of having an LLM run on a server somewhere in the cloud is quickly becoming outmoded.

Binh Pham experimented with a Raspberry Pi Zero, effectively turning the device into a small USB drive that can run an LLM locally with no extras needed. The project was largely facilitated thanks to llama.cpp and llamafile, a combination of an instruction set and a series


Open Source LLM

9to5Neural: xAI unveiling Grok 3 tonight — could GPT-4.5 steal the show?

Welcome to 9to5Neural . AI moves fast. We help you keep up. OpenAI says GPT-4.5 is coming to ChatGPT in a matter of weeks. But first, xAI will unveil Grok 3 tonight. Now the countdown is on for OpenAI to do the funniest thing ever…

Fresh off the heels of threatening to buy OpenAI, Elon Musk announced on Saturday that Grok 3 will arrive tonight in the form of a live demo scheduled for 8 p.m. PT.

Musk hypes up Grok 3

Menu