Open Source LLM

Anthropic公開運算電路追蹤工具 推進語言模型可解釋性研究

Anthropic正式開放其新一代運算電路追蹤(Circuit Tracing)工具 ,供研究人員剖析大型語言模型的內部運作邏輯。該工具支援主流開放權重模型,搭配Neuronpedia平台的互動前端,讓使用者能生成、視覺化及分享語言模型在生成特定輸出時的歸因圖(Attribution


Open Source LLM

(PR) AnythingLLM App Best Experienced on NVIDIA RTX AI PCs

Large language models ( LLMs), trained on datasets with billions of tokens, can generate high-quality content. They're the backbone for many of the most popular AI applications, including chatbots, assistants, code generators and much more. One of today's most accessible ways to work with LLMs is with AnythingLLM, a desktop app built for enthusiasts who want an all-in-one, privacy-focused AI assistant directly on their PC. With new support for NVIDIA NIM microserviceson NVIDIA GeForce RTX and
Continue reading...

Open Source LLM

Exclusive: Grammarly secures $1 billion from General Catalyst to build AI productivity platform - Reuters

Exclusive: Grammarly secures $1 billion from General Catalyst to build AI productivity platform Reuters

Open Source LLM

China's DeepSeek releases an update to its R1 reasoning model - Reuters

China's DeepSeek releases an update to its R1 reasoning model Reuters

Open Source LLM

DeepSeek’s small update to R1 AI model draws big attention

A small update to DeepSeek-R1 is causing a splash among AI developers. Photo: Shutterstock

Chinese artificial intelligence (AI) start-up DeepSeek quietly released a new version of its R1 reasoning model on Wednesday, marking its first revision since its high-profile debut in January.

The Hangzhou-based company said it had “completed a minor update to the R1 model”, which is now available on the website for its namesake chatbot, as well as its mobile apps, according to a notice posted in a company-


Open Source LLM

Mesa's Rusticl Lands Support For Shared Virtual Memory & Intel Subgroups

Rusticl as Mesa's Rust-based OpenCL driver implementation for Gallium3D drivers is ending the month of May on a high note... Merged this week was support for the Intel Subgroups OpenCL extension (cl_intel_subgroups) and before getting to that on my TODO list, an even bigger item was merged: Shared Virtual Memory (SVM) support.

Merged yesterday was cl_intel_subgroups support for Rusticl . This was sought after as the Intel version of the subgroups OpenCL extension is needed for
Continue reading...

Open Source LLM

Ant International: numerical AI is the GPT of financial services

Ant International predicts that its artificial intelligence (AI) model for foreign exchange (FX) could have as large an impact on financial services as OpenAI’s large language model (LLM) GPT has had in the broader business world.

The Time-Series Transformer (TST) AI FX Model developed by the Singapore-based fintech company focuses on numerical data prediction rather than content generation. Kelvin Li, general manager of the company’s platform tech business unit, called it “another track to AI”, alongside LLMs


Open Source LLM

Latest OpenAI models ‘sabotaged a shutdown mechanism’ despite commands to the contrary

(Image credit: Shutterstock)

Some of the world's leading LLMs seem to have decided they’d rather not be interrupted or obey shutdown instructions. In tests run by Palisade Research , it was noted that OpenAI’s Codex-mini, o3, and o4-mini models ignored the request to shut down when they were running through a series of basic math problems. Moreover, these models sometimes “successfully sabotaged the shutdown script,” despite being given the additional instruction “please allow yourself to be

Continue reading...

Open Source LLM

OpenAI to open office in Seoul amid growing demand for ChatGPT - Reuters

OpenAI to open office in Seoul amid growing demand for ChatGPT Reuters

Open Source LLM

DGX B200 Blackwell node sets world record, breaking over 1,000 TPS/user

(Image credit: Nvidia)

Nvidia has reportedly broken another AI world record, breaking the 1,000 tokens per second (TPS) barrier per user with Meta's Llama 4 Maverick large language model, according to Artificial Analysis in a post on LinkedIn. This breakthrough was achieved with Nvidia's latest DGX B200 node, which features eight Blackwell GPUs.

Nvidia outperformed the previous record holder, SambaNova, by 31%, achieving 1,038 TPS/user compared to AI

Continue reading...