Open Source LLM

MIT 揭 LLM 捉錯用神 只認句式唔認字 黑客輕易繞過防護機制

MIT 最新研究發現,大型語言模型(LLM)回答問題時有時會「學錯重點」,依賴訓練期間學到的語法模式作答,而非真正理解問題內容。這種現象會令模型處理新任務時出現意外失誤,影響客戶查詢處理、臨床記錄摘要及財務報告


Open Source LLM

AMD ROCm 7.1.1 Released With RHEL 10.1 Support, More Models Working On RDNA4

Following the release of ROCm 7.1 from just under one month ago, ROCm 7.1.1 is now available with expanded Linux operating system support, continued Instinct MI350 series work, more large language models working on RDNA4 GPUs, and other enhancements.

ROCm 7.1.1 is now available as the newest point release for this open-source AMD GPU compute stack for Radeon and Instinct hardware. Some of the ROCm 7.1.1 release highlights include:

- Support for Red
Continue reading...

Open Source LLM

Researchers discover a shortcoming that makes LLMs less reliable

Large language models can learn to mistakenly link certain sentence patterns with specific topics — and may then repeat these patterns instead of reasoning.

Open Source LLM

OpenAI projects 220 million paying ChatGPT users by 2030, The Information Reports - Reuters

OpenAI projects 220 million paying ChatGPT users by 2030, The Information Reports Reuters

Open Source LLM

Intel LLM Scaler vLLM Update Supports More Models

Intel software engineers continue to be hard at work on LLM-Scaler as their solution for running vLLM on Intel GPUs in a Docker containerized environment . A new beta release of LLM-Scaler built around vLLM was released overnight with support for running more large language models.

Since the "LLM-Scaler 1.0" debut of the project back in August there have been frequent updates for expanding LLM coverage on Intel GPUs and exposing more features for harnessing the AI compute power on Intel graphics hardware. The versioning scheme though remains
Continue reading...

Open Source LLM

新加坡國家AI計劃放棄Meta模型 轉向阿里旗下通義千問

據內媒報道,新加坡國家人工智能計劃(AISG)正進行一次重大戰略調整,在其最新東南亞語言大模型項目中放棄Meta模型,轉向阿里(9988)旗下通義千問(Qwen)開源架構,標誌着中國開源AI模型在全球影響力版圖中的一次關


Open Source LLM

Anthropic’s new model is its latest frontier in the AI agent battle — but it’s still facing cybersecurity concerns

The AI labs never sleep — especially the week before Thanksgiving, it seems. Days after Google’s buzzworthy Gemini 3, and OpenAI’s updated agentic coding model, Anthropic has announced Claude Opus 4.5, which it bills as “the best model in the world for coding, agents, and computer use,” claiming it has leapfrogged even Gemini 3 in […]

Open Source LLM

Anthropic reveals new Opus 4.5 model, brings Claude Code to the Mac app

Anthropic has announced its latest AI model with Claude Opus 4.5. The company has also expanded Claude Code availability to the Claude desktop app for the first time.

Anthropic describes its new Opus 4.5 model as “intelligent, efficient, and the best model in the world for coding, agents, and computer use.” It follows Opus 4.1, which Anthropic released in August. The company shares internal first impressions of new model:

As our Anthropic colleagues tested the model before release, we heard remarkably


Open Source LLM

(PR) AMD MI300X Powers Training of Zyphra's First Large-Scale MoE Model, ZAYA1

AMD (NASDAQ: AMD) announced that Zyphra has achieved a major milestone in large-scale AI model training with the development of ZAYA1, the first large-scale Mixture-of-Experts (MoE) foundation model trained using an AMD GPU and networking platform. Using AMD Instinct MI300X GPUs and AMD Pensando networking and enabled by the AMD ROCm open software stack, the achievement is detailed in a Zyphra technical report published today.

Results from Zyphra show that the model delivers competitive or superior performance to
Continue reading...

Open Source LLM

New Apple study shows LLMs can tell what you’re doing from audio and motion data

Apple researchers have published a study that looks into how LLMs can analyze audio and motion data to get a better overview of the user’s activities. Here are the details.

A new paper titled “ ” offers insight into how Apple may be considering incorporating LLM analysis alongside traditional sensor data to gain a more precise understanding of user activity.

This, they argue, has great potential to make activity analysis more precise, even in situations where there isn’t enough sensor data.

From the researchers:

“Sensor data streams

Continue reading...
Menu