Open Source LLM

PoCL 6.0 OpenCL Implementation Brings OpenMP For CPU Driver, More Remote Driver Features

The Portable Computing Language "PoCL" that started off as a CPU-based OpenCL implementation has grown to support multiple hardware targets from NVIDIA PTX to Intel Level Zero to AMD ROCm and other innovations like a recent remote driver for transparent OpenCL across networked systems . PoCL 6.0 was released today for delivering the latest enhancements to this independent OpenCL compute implementation and continuing to enhance support for its different hardware targets.

With PoCL 6.0 for its remote OpenCL driver there is now support for coarse-grained Shared Virtual Memory (
Continue reading...

Open Source LLM|Intel

Intel Releases OpenVINO 2024.2 With Llama 3 Optimizations, More AVX2 & AVX-512 Optimizations

Intel today released OpenVINO 2024.2, the newest version of its open-source AI toolkit for optimizing and deploying deep learning (A) inference models across a range of AI frameworks and broad hardware types.

With OpenVINO 2024.2 they have continued optimizing for Meta's Llama 3 large language model. OpenVINO 2024.2 brings more Llama 3 optimizations for execution across CPUs, integrated GPUs, and discrete GPUs to further enhance performance while yielding more efficient memory use too.
Continue reading...

Open Source LLM

Nvidia開源Nemotron-4 340B家族,以供開發者建置大型語言模型

Nvidia上週 開源了Nemotron-4 340B模型家族 ,它包含了基礎模型、指令模型及獎勵模型,可用來生成合成資料,藉以訓練大型語言模型(LLM),現已可 自Hugging Face下載 ,之後也能透過Nvidia網站以API及NIM微服務來存取模型。

Nvidia表示,高品質的


Open Source LLM

(PR) NVIDIA MLPerf Training Results Showcase Unprecedented Performance and Elasticity

The full-stack NVIDIA accelerated computing platform has once again demonstrated exceptional performance in the latest MLPerf Training v4.0 benchmarks. NVIDIA more than tripled the performance on the large language model (LLM) benchmark, based on GPT-3 175B, compared to the record-setting NVIDIA submission made last year. Using an AI supercomputer featuring 11,616 NVIDIA H100 Tensor Core GPUs connected with NVIDIA Quantum-2 InfiniBand networking, NVIDIA achieved this remarkable feat through larger scale -

Open Source LLM|Intel

(PR) Intel Claims Lower Cost Alternative for AI Compute and GenAI With Latest Gaudi 2 MLCommons Results

Today, MLCommons published results of its industry AI performance benchmark, MLPerf Training v4.0. Intel's results demonstrate the choice that Intel Gaudi 2 AI accelerators give enterprises and customers. Community-based software simplifies generative AI (GenAI) development and industry-standard Ethernet networking enables flexible scaling of AI systems. For the first time on the MLPerf benchmark, Intel submitted results on a large Gaudi 2 system (1,024 Gaudi 2 accelerators) trained in Intel Tiber Developer Cloud to demonstrate Gaudi 2 performance

Open Source LLM|Intel

(PR) Intel Submits Gaudi 2 Results on MLCommons' Newest Benchmark

Today, MLCommons published results of its industry AI performance benchmark, MLPerf Training v4.0. Intel's results demonstrate the choice that Intel Gaudi 2 AI accelerators give enterprises and customers. Community-based software simplifies generative AI (GenAI) development and industry-standard Ethernet networking enables flexible scaling of AI systems. For the first time on the MLPerf benchmark, Intel submitted results on a large Gaudi 2 system (1,024 Gaudi 2 accelerators) trained in Intel Tiber Developer Cloud to demonstrate Gaudi 2 performance

Open Source LLM|Intel

Intel's oneDNN 3.5 Begins Optimizing For Xe2, More Xeon 6 Tuning

Intel's oneDNN 3.5 has been released as this Deep Neural Network Library for the oneAPI specification and now part of the UXL Foundation. With oneDNN 3.5 comes more performance optimizations for existing and upcoming Intel hardware.

The oneDNN 3.5 release has improved performance for 4th Gen Xeon Scalable "Sapphire Rapids" processors and improving the performance for Xeon 6 with the recently launched SIerra Forest and the upcoming Granite Rapids CPUs. The oneDNN 3.5 release also has common tuning for enhancing the performance of the
Continue reading...

Open Source LLM

Mold 2.32 Released With Increased LLVM LLD Compatibility, Faster Identical Code Folding

Mold 2.32 is out as the newest feature release for this high speed code linker that rivals LLVM LLD and GNU Gold.

With Mold 2.32 comes support for faster Identical Code Folding (ICF) as a means of finding identical functions and merging them to reduce the size of the output file.Identical Code Folding with Mold has been found to be very helpful for template-heavy C++ programs. With Mold 2.32, their ICF algorithm is around 50% faster than with
Continue reading...

Open Source LLM

Stability AI釋出文字生成聲音模型開源版本Stable Audio Open

Stability AI週三(6/5)釋出了文字生成聲音模型的開源版本 Stable Audio Open ,在使用者輸入文字描述後,它便能生成長達47秒的樣本與聲音效果。

Stability AI以超過48萬個聲音紀錄來訓練Stable Audio Open模型,其中超過9成的紀錄來自Freesound,另有少數來


Open Source LLM

AMD ROCm 6.1.2 Released With Fixes & Optimizations

ROCm 6.1.2 is out today as the newest update to AMD's open-source GPU compute stack for Linux systems and with growing support for Windows Subsystem for Linux.

Being a point release, ROCm 6.1.2 is mostly focused on providing various fixes and optimizations. Arguably most significant is having preview support for the upcoming Ubuntu 22.04.5 LTS point release with both the Linux 5.15 GA kernel and the Linux 6.8 HWE kernel. Ubuntu 22
Continue reading...