<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://hpcgroup.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://hpcgroup.github.io/" rel="alternate" type="text/html" /><updated>2026-06-27T04:43:24+00:00</updated><id>https://hpcgroup.github.io/feed.xml</id><title type="html">Parallel Software and Systems Group</title><subtitle>Website for PSSG.</subtitle><author><name>UMD HPC Group</name></author><entry><title type="html">PSSG at ACM/IEEE Supercomputing 2025</title><link href="https://hpcgroup.github.io/blog/2025/sc25/" rel="alternate" type="text/html" title="PSSG at ACM/IEEE Supercomputing 2025" /><published>2025-11-10T00:00:00+00:00</published><updated>2025-11-10T00:00:00+00:00</updated><id>https://hpcgroup.github.io/blog/2025/sc25</id><content type="html" xml:base="https://hpcgroup.github.io/blog/2025/sc25/"><![CDATA[<p>PSSG will have a strong presence at SC ‘25, the premier international conference for high-performance computing, networking, storage, and analysis. We are excited to be contributing in several impactful ways.</p>

<h3 id="technical-paper-presentations">Technical Paper Presentations</h3>
<p><a href="https://linktr.ee/adityaranjan">Aditya Ranjan</a> (graduated MS student) and team will be presenting their work on <a href="https://github.com/hpcgroup/plexus">Plexus</a>: a highly scalable GNN training framework that achieves unprecedented speedups of 2.3-12.5X over prior state of the art, and a reduction in time-to-solution by 5.2-8.7X on NERSC’s Perlmutter (scaled up to 2,048 GPUs) and 7.0-54.2X on Frontier (scaled up to 1,024 GPUs).
<a href="https://sc25.conference-program.com/presentation/?id=pap882&amp;sess=sess282">[Talk Link]</a></p>

<h3 id="posters">Posters</h3>
<p>Prajwal Singhania (third-year PhD student) and Lannie Dalton Hough (second-year PhD student) will be presenting their poster titled “Understanding Communication Bottlenecks in Multi-Node LLM Inference”, 
about investigating the scaling bottlenecks in multi-node LLM inference using <a href="https://github.com/axonn-ai/yalis">YALIS</a>, our prototype inference engine.</p>

<p>Onur Cankur (final-year PhD student) will be presenting his poster titled “Understanding GPU Utilization Using LDMS Data on Perlmutter”
about analyzing spatial and temporal trends of workloads on Perlmutter to reveal inefficiencies and imbalances that can guide workload optimization.</p>

<p>Cunyang Wei (second-year PhD student) will be presenting his poster titled “Unmasking Performance Variability in GPU Codes on Supercomputers”. The poster is about evaluating and predicting performance variability on GPU supercomputing systems (Perlmutter and Frontier). It also provides several recommendations for system administrators and users to mitigate performance variability.</p>

<p>Keshav Pradeep (senior undergraduate student) will be presenting a poster titled “Optimizing Collectives with Large Payloads on GPU-based Supercomputers.” The poster evaluates the performance of existing collective libraries (RCCL and Cray-MPICH) on Frontier, identifying key bottlenecks that limit scalability in large model training. It introduces PCCL, which scales efficiently to 2048 GPUs, achieving up to 33× faster all-gather operations and nearly 5× end-to-end training speedups.</p>

<h3 id="tutorials">Tutorials</h3>
<p>The AxoNN and YALIS team (Prajwal Singhania, Lannie Dalton Hough, Cunyang Wei) will present a tutorial on <a href="https://github.com/axonn-ai/axonn">AxoNN</a>, a highly scalable and easy-to-use parallel framework for AI training. Attendees of the tutorial will learn how to use state-of-the-art parallel algorithms in AxoNN to parallelize their GPU training workloads with minimal code changes. We will demonstrate this using a very popular use case—finetuning open-source LLMs like Llama-3 on instruction finetuning data. We will also dive into the basics of LLM inference with <a href="https://github.com/axonn-ai/yalis">YALIS</a> and vLLM.</p>

<h3 id="umd-booth-talks">UMD Booth Talks</h3>
<p>PhD students —  <a href="https://joy-kitson.github.io/">Joy Kitson</a>, <a href="https://jhdavis8.github.io/">Josh Davis</a>, <a href="https://prajwal1210.github.io/">Prajwal Singhania</a>, 
<a href="https://cunyangwei.github.io/">Cunyang Wei</a>, <a href="https://ldhough.github.io/">Lannie Dalton Hough</a>, Emir Gencer, and Klaudiusz Rydzy — will be presenting their research work at the UMD Exhibitor Booth at SC25. A schedule of the talks can be found below.</p>

<ul>
  <li>Tuesday, November 18
    <ul>
      <li>12:10 p.m. “Understanding and Mitigating Communication Bottlenecks in Multi-Node LLM Inference” – Prajwal Singhania</li>
      <li>1:10 p.m. “Unmasking Performance Variability in GPU Codes on Production Supercomputers” – Cunyang Wei</li>
      <li>3:10 p.m. “Layer-Aware Asymmetric Tensor Parallelism for Heterogeneous LLM Inference” – Lannie Dalton Hough</li>
    </ul>
  </li>
  <li>Wednesday, November 19
    <ul>
      <li>10:10 a.m. “Understanding GPU Utilization Using LDMS Data on Perlmutter” – Onur Cankur</li>
      <li>12:10 p.m. “Modeling Interdependent Social and Biological Contagions on Massive Multi-layer Networks” – Joy Kitson</li>
      <li>1:10 p.m. “Can LLMs Explain GPU Kernel Performance?” – Joshua H. Davis</li>
    </ul>
  </li>
  <li>Thursday, November 20
    <ul>
      <li>10:10 a.m. “ucTrace: Monitoring MPI Communication with UCX Tracing” – Emir Gencer</li>
      <li>12:10 p.m. “Can Agentic Frameworks Identify and Fix Performance Issues” – Klaudiusz Rydzy</li>
    </ul>
  </li>
</ul>

<hr />

<p>We look forward to connecting with colleagues and collaborators at SC ‘25. Stay tuned for more updates as the conference approaches!</p>]]></content><author><name>UMD HPC Group</name></author><category term="General" /><category term="conferences" /><summary type="html"><![CDATA[PSSG will have a strong presence at SC ‘25, the premier international conference for high-performance computing, networking, storage, and analysis. We are excited to be contributing in several impactful ways.]]></summary></entry><entry><title type="html">Beyond NCCL: Faster Inter-Node All-Reduce for Decode-Heavy LLM Inference</title><link href="https://hpcgroup.github.io/blog/2025/beyond-nccl/" rel="alternate" type="text/html" title="Beyond NCCL: Faster Inter-Node All-Reduce for Decode-Heavy LLM Inference" /><published>2025-09-16T00:00:00+00:00</published><updated>2025-09-16T00:00:00+00:00</updated><id>https://hpcgroup.github.io/blog/2025/beyond-nccl</id><content type="html" xml:base="https://hpcgroup.github.io/blog/2025/beyond-nccl/"><![CDATA[<h2 id="introduction">Introduction</h2>

<p>As large language models grow in size, a single node —
typically 4-8 NVLink-connected GPUs — is no longer sufficent to hold the
model and run inference.
Models like
Llama-405B or Nemotron-Ultra-253B can require 16–32 GPUs even for modest batch
sizes. On most HPC systems, that means stepping outside the comfort of NVLink
into the slower world of inter-node networking, where communication costs
dominate.</p>

<p>So what’s the best way to run inference in this multi-node world?</p>

<p>For moderate context lengths (a few thousand tokens), there are two main
distributed inference strategies:</p>

<p><strong>Tensor Parallelism (TP)</strong> splits each layer’s weights across
  GPUs and requires an All-Reduce collective after each matrix multiply.</p>

<p><strong>Pipeline Parallelism (PP)</strong> splits the model into contiguous subsets of layers.
	It requires only point-to-point communication but introduces sequential dependencies and pipeline “bubbles.”</p>

<p>Conventional wisdom says that TP within a node and PP across nodes is the best strategy,
as it reduces the inter-node message volume. But on HPC systems, with advanced
interconnect fabrics like Slingshot or InfiniBand, the trade-off isn’t that simple:</p>

<ul>
  <li>PP reduces inter-node message volume, but also introduces sequential
dependencies and pipeline bubbles. This shows up in single-batch offline
inference scenarios where pipeline latency scales with number of nodes.</li>
  <li>TP has a high message volume, but also parallelizes all the matrix multiplications.</li>
</ul>

<h2 id="llm-inference-beyond-a-single-node">LLM Inference Beyond a Single Node</h2>
<p>To understand how LLM frameworks scale beyond a single node in <strong><em>offline single-batch inference scenarios</em></strong>, 
we built <a href="https://github.com/axonn-ai/yalis">YALIS</a> (Yet Another LLM Inference System), a light-weight and performant
inference engine prototype that gives us full instrumentation flexibility. We
validate its performance against vLLM and run strong scaling experiments for
both frameworks on the Perlmutter supercomputer (4X NVIDIA 80GB A100s per node
connected via NVLink; Slingshot-11 interconnect between nodes). Our experiment 
here uses the Llama-3.1-70B-Instruct model with a batch size of 16, prefill
length of 2480, and decode length of 2048.</p>

<figure>
  <a href="/assets/images/posts/yalis_vllm_scaling_pm.png">
    <img src="/assets/images/posts/yalis_vllm_scaling_pm.png" alt="Strong scaling results for YALIS and vLLM" width="800" />
  </a>
  <figcaption>End-to-end latency when strong scaling Llama3.1-70B from 1 to 8 nodes (4 to 32 GPUs).</figcaption>
</figure>

<p><em>Findings</em>: Using the common wisdom TP+PP approach, end-to-end latency increases
with increasing GPU/node count. Remember that under ideal strong scaling (fixed total
problem size with increasing GPU count), latency should decrease with
increasing resources. None of the parallelism strategies achieve that, with
TP+PP being the worst. We believe this is due to the fact that PP is suited more
to online high throughput scenarios where a constant influx of batches can hide
pipeline bubbles and hide the sequential dependencies.</p>

<h2 id="a-deeper-look-with-yalis">A Deeper Look with YALIS</h2>
<p>Noting that YALIS performs similar to vLLM, we instrument YALIS to get detailed
breakdowns and understand what is going on.
We study separately the two phases of LLM inference: <em>prefill</em> and <em>decode</em>. 
We subcategorize the time spent in each of these phases into: Matmul (Time spent in matrix multiplications),
Communication, FlashAttn (Time spent in FlashAttention), Other (Time spent in
all other GPU operations), Idle (GPU Idle Time).</p>

<figure>
  <a href="/assets/images/posts/breakdown_all_pm.png">
    <img src="/assets/images/posts/breakdown_all_pm.png" alt="YALIS Inference Time Breakdown" width="800" />
  </a>
  <figcaption>Prefill and decode (×10 steps) time breakdown for Llama-3.1-70B Instruct when scaling YALIS with TP within and across nodes.</figcaption>
</figure>

<p><em>Findings</em>: The compute time (Matmul + FlashAttn + Other) decreases as expected, demonstrating ideal strong scaling with increasing GPU count. However, the communication
time blows up substantially going from single to multi-node. At 32 GPUs, it is
&gt;50% of the time for both prefill and decode.</p>

<p>This shows that communication is a big bottleneck for both prefill and decode.
While both phases do All-Reduce communication among the GPUs, there is a
key difference in the two phases with respect to the 
<strong><em>All-Reduce message sizes</em></strong>. In this particular example, the prefill All-Reduce 
message size is 620 MiB (very large message regime), whereas the decode
All-Reduce message size is 256 KiB (small message regime).</p>

<h2 id="beyond-nccl--tailoring-all-reduce-for-decode">Beyond NCCL — Tailoring All-Reduce for Decode</h2>
<p>Next, we turned our focus to communication in the decode phase. Although prefill
also suffers from communication overhead, decode is the bigger concern for
generation-heavy workloads: its auto-regressive token-by-token nature — issuing
many small, latency-sensitive messages — dominates end-to-end latency.</p>

<p>Most inference frameworks rely on NCCL (on NVIDIA systems) for collective
communication. While NCCL excels within a node for small messages, its
performance drops once communication crosses nodes. On non-NVLink interconnects
(without NVSwitch), NCCL typically uses either Ring All-Reduce or Tree +
Broadcast All-Reduce. But are these algorithms optimal for the small-message,
low-latency regime of decode? Other collective libraries, such as MPI, offer a
broader range of algorithms — including recursive approaches that are
theoretically better here — but lack features important for inference, like CUDA
graph friendliness and low kernel launch overhead.</p>

<p>Seeing this gap, we set out to design something better. Our answer - 
<strong><a href="https://github.com/pssg-int/nvshmem-allreduce/">NVRAR</a></strong>: <em>a GPU-initiated <u>NV</u>SHMEM <u>R</u>ecursive <u>A</u>ll-<u>R</u>educe implementation</em> , 
tuned for decode’s small message sizes beyond a single node, and built to
integrate cleanly with CUDA graphs and low-overhead inference. NVRAR uses a
heirarchical approach, leveraging NCCL for intra-node Reduce-Scatter and
All-Gather and a custom inter-node Recursive All-Reduce implementation with low
synchronization overhead.</p>

<p>Early experiments with NVRAR show consistent speedups over NCCL in the decode message size regime.</p>

<figure>
  <a href="/assets/images/posts/heatmap_nccl_nvrar_pm.png">
    <img src="/assets/images/posts/heatmap_nccl_nvrar_pm.png" alt="Speedup of NVRAR over NCCL" width="800" />
  </a>
  <figcaption>Speedup of NVRAR over NCCL (no CUDA Graphs, 200 warmup iterations + 1000 timed iterations).</figcaption>
</figure>

<figure>
  <a href="/assets/images/posts/breakdown_Decode_nvshmem_vs_nccl.png">
    <img src="/assets/images/posts/breakdown_Decode_nvshmem_vs_nccl.png" alt="Decode Latency Breakdown NCCL vs NVRAR" width="800" />
  </a>
  <figcaption>Comparing the breakdown of 10 Decode steps with NCCL and NVRAR for Llama-3.1-70B and 405B Instruct 
   models at different GPU counts.</figcaption>
</figure>

<p><em>Findings</em>: NVRAR achieves modest speedups compared to NCCL in the 256 KiB to
2 MiB message size range. Further, when plugged into YALIS, we see a reduction in
decode latencies (TBT times) by up to <strong>15%</strong> for Llama-3.1-70B Instruct on 16 GPUs,
and <strong>38.8%</strong> for Llama-3.1-405B-Instruct on 64 GPUs.</p>

<p>We’re also making both <a href="https://github.com/axonn-ai/yalis">YALIS</a> and <a href="https://github.com/hpcgroup/nvrar">NVRAR</a> open-source so that others can
experiment, benchmark, and provide feedback. The full design details and extended evaluation can found in our <a href="https://arxiv.org/abs/2511.09557">paper</a>.</p>

<h2 id="references-and-links">References and Links</h2>

<ul>
  <li>YALIS: <a href="https://github.com/axonn-ai/yalis">https://github.com/axonn-ai/yalis</a></li>
  <li>vLLM: <a href="https://github.com/vllm-project/vllm">https://github.com/vllm-project/vllm</a></li>
  <li>NCCL: <a href="https://developer.nvidia.com/nccl">https://developer.nvidia.com/nccl</a></li>
  <li>NVSHMEM: <a href="https://developer.nvidia.com/nvshmem">https://developer.nvidia.com/nvshmem</a></li>
  <li>NVRAR: <a href="https://github.com/hpcgroup/nvrar">https://github.com/hpcgroup/nvrar</a></li>
</ul>]]></content><author><name>UMD HPC Group</name></author><category term="HPC for ML" /><category term="LLM" /><category term="collectives" /><summary type="html"><![CDATA[Introduction]]></summary></entry><entry><title type="html">PSSG at ACM/IEEE Supercomputing 2024</title><link href="https://hpcgroup.github.io/blog/2024/sc24/" rel="alternate" type="text/html" title="PSSG at ACM/IEEE Supercomputing 2024" /><published>2024-10-16T00:00:00+00:00</published><updated>2024-10-16T00:00:00+00:00</updated><id>https://hpcgroup.github.io/blog/2024/sc24</id><content type="html" xml:base="https://hpcgroup.github.io/blog/2024/sc24/"><![CDATA[<p>PSSG will have a strong presence at SC ‘24, the premier international conference for high-performance computing, networking, storage, and analysis. We are excited to be contributing in several impactful ways.</p>

<h3 id="technical-paper-presentations">Technical Paper Presentations</h3>
<p><a href="https://dando18.github.io/">Daniel Nichols</a> (final-year PhD student) will be presenting his paper on a probabilistic approach to selecting build configurations in package managers using historical build data and integrating this approach into the Spack package manager using probabilistic answer set programming.
<a href="https://sc24.conference-program.com/presentation/?id=pap175&amp;sess=sess390">[Talk Link]</a></p>

<p><a href="https://siddharth9820.github.io/">Siddharth Singh</a> (final-year PhD student) and his team will be presenting their work on scaling <a href="https://github.com/axonn-ai/axonn">AxoNN</a>, a 4D parallel AI training framework, on the Frontier, Perlmutter, and Alps supercomputers. They achieve an impressive 1.423 Exaflop/s on 6,144 NVIDIA H100 GPUs, 1.381 Exaflop/s on 32,768 AMD MI250X GCDs, and 620.1 Petaflop/s on 4,096 NVIDIA A100 GPUs in half-precision (bf16) training of LLMs. This effort allows them to analyze the side effect of scaling AI models via experiments that explore <em>catastrophic memorization</em>, where models are sufficiently large to memorize training data in a single pass, and present a preventative approach. <a href="https://sc24.conference-program.com/presentation/?id=gb102&amp;sess=sess496">[Talk Link]</a></p>

<h3 id="awards">Awards</h3>
<p>Daniel Nichols is a recipient of the <a href="https://awards.acm.org/hpc-fellows">2024 ACM-IEEE CS George Michael Memorial HPC Fellowship</a>. This award will be presented during the awards ceremony SC ‘24.</p>

<p>The AxoNN team’s submission has been selected as a finalist for the prestigious <a href="https://awards.acm.org/bell">ACM Gordon Bell Prize</a> at SC ‘24. <a href="https://sc24.conference-program.com/presentation/?id=gb102&amp;sess=sess496">[Talk Link]</a></p>

<h3 id="posters">Posters</h3>
<p>Aman Chaturvedi (undergraduate student) will be presenting his poster titled “Creating Code LLMs for HPC: It’s LLMs All the Way Down,” about creating <a href="https://huggingface.co/hpcgroup/hpc-coder-v2-6.7b">HPC-Coder-v2</a> with synthetic code data. It’s shown to be the best LLM at writing parallel code with less than 30B parameters.</p>

<p>Aditya Tomar (UC Berkeley undergraduate student) will be presenting his poster titled “Eve: Less Memory, Same Might,” about building Eve, an approximation of the AdamW optimizer for LLM training, which uses nearly 25% less memory while providing the same convergence as AdamW.</p>

<h3 id="tutorials">Tutorials</h3>
<p>The AxoNN team, led by Siddharth Singh, will present a tutorial on <a href="https://github.com/axonn-ai/axonn">AxoNN</a>, a highly scalable and easy-to-use parallel framework for AI training. Attendees of the tutorial will learn how to use state-of-the-art parallel algorithms in AxoNN to parallelize their GPU training workloads with minimal code changes. We will demonstrate this using a very popular use case—finetuning open-source LLMs like Llama-3 on instruction finetuning data.</p>

<h3 id="birds-of-a-feather-session">Birds of a Feather Session</h3>
<p>Daniel Nichols and others will organize a BoF session titled “Toward Integrating LLMs in HPC Software Development.” LLM-based coding assistants have already proven to be useful tools for aiding software developers, and adapting these tools in HPC software development will greatly improve the quality and time-to-development of scientific codes. This will create an environment where researchers can devote more attention to scientific challenges and less to software development intricacies, driving scientific progress forward. This BoF will provide a place for the community to discuss the use of LLMs for HPC software development. <a href="https://sc24.conference-program.com/presentation/?id=bof229&amp;sess=sess662">[Session Link]</a> <a href="https://parallelcodefoundry.github.io/sc24-bof.html">[BoF Website]</a></p>

<h3 id="umd-booth-talks">UMD Booth Talks</h3>
<p>PhD students — Daniel Nichols, Siddharth Singh, <a href="https://joy-kitson.github.io/">Joy Kitson</a>, <a href="https://jhdavis8.github.io/">Josh Davis</a>, and <a href="https://prajwal1210.github.io/">Prajwal Singhania</a>—will be presenting their research work at the UMD Exhibitor Booth at SC24. A schedule of the talks can be found below.</p>

<ul>
  <li>Tuesday, November 19
    <ul>
      <li>10:10 a.m. “Taking GPU Programming Models to Task for Performance Portability” – Joshua Davis</li>
      <li>1:10 p.m. “Eve: Pruning Adam’s State for Scalable Deep Learning” – Aditya Tomar</li>
      <li>3:10 p.m. “Strategies for Parallelising an Agent-Based Model of Infectious Disease Spread” – Joy Kitson</li>
    </ul>
  </li>
  <li>Wednesday, November 20
    <ul>
      <li>10:10 a.m. “Loki: Low-rank Keys for Efficient Sparse Attention” – Prajwal Singhania</li>
      <li>1:10 p.m. “Improving Build Likelihood in Package Managers with Probabilistic Constraints” – Daniel Nichols</li>
      <li>3:10 p.m. “Insights from Longitudinal GPU Workload Monitoring on Perlmutter” – Onur Cankur</li>
    </ul>
  </li>
  <li>Thursday, November 21
    <ul>
      <li>10:10 a.m. “Creating Code LLMs for HPC: It’s LLMs All the Way Down” – Aman Chaturvedi</li>
      <li>1:10 p.m. “A hybrid tensor-expert-data parallelism approach to optimize mixture-of-experts training” – Siddharth Singh</li>
    </ul>
  </li>
</ul>

<hr />

<p>We look forward to connecting with colleagues and collaborators at SC ‘24. Stay tuned for more updates as the conference approaches!</p>]]></content><author><name>UMD HPC Group</name></author><category term="General" /><category term="conferences" /><summary type="html"><![CDATA[PSSG will have a strong presence at SC ‘24, the premier international conference for high-performance computing, networking, storage, and analysis. We are excited to be contributing in several impactful ways.]]></summary></entry><entry><title type="html">ParEval Leaderboard: Evaluating the Ability of Large Language Models to Generate Parallel Code</title><link href="https://hpcgroup.github.io/blog/2024/pareval/" rel="alternate" type="text/html" title="ParEval Leaderboard: Evaluating the Ability of Large Language Models to Generate Parallel Code" /><published>2024-02-07T00:00:00+00:00</published><updated>2024-02-07T00:00:00+00:00</updated><id>https://hpcgroup.github.io/blog/2024/pareval</id><content type="html" xml:base="https://hpcgroup.github.io/blog/2024/pareval/"><![CDATA[<p>We introduced the ParEval benchmark in “<a href="https://arxiv.org/abs/2401.12554">Can Large Language Models Write
Parallel Code?</a>” to evaluate the capability of
LLMs at parallel code generation. We found a significant gap between their
ability to generate sequential vs parallel code for a large array of
computational problems and parallel programming models. On this page we keep an
up-to-date table tracking the progress of state-of-the-art LLMs on ParEval.</p>

<h2 id="pareval-results">ParEval Results</h2>

<table>
  <thead>
    <tr>
      <th>Model</th>
      <th>No. Parameters</th>
      <th>HumanEval<br />pass@1</th>
      <th>ParEval Serial<br />pass@1</th>
      <th>ParEval Parallel<br />pass@1</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><a href="https://huggingface.co/bigcode/starcoder2-3b"><i class="fas fa-fw fa-solid fa-link"></i></a> StarCoder2-3B</td>
      <td>3B</td>
      <td>31.7</td>
      <td>42.7</td>
      <td>9.6</td>
    </tr>
    <tr>
      <td><a href="https://huggingface.co/bigcode/starcoder2-7b"><i class="fas fa-fw fa-solid fa-link"></i></a> StarCoder2-7B</td>
      <td>7B</td>
      <td>35.4</td>
      <td>59.4</td>
      <td>15.9</td>
    </tr>
    <tr>
      <td><a href="https://huggingface.co/codellama/CodeLlama-7b-hf"><i class="fas fa-fw fa-solid fa-link"></i></a> CodeLlama-7B</td>
      <td>7B</td>
      <td>29.9</td>
      <td>48.4</td>
      <td>15.3</td>
    </tr>
    <tr>
      <td><a href="https://huggingface.co/codellama/CodeLlama-13b-hf"><i class="fas fa-fw fa-solid fa-link"></i></a> CodeLlama-13B</td>
      <td>13B</td>
      <td>35.0</td>
      <td>52.8</td>
      <td>17.4</td>
    </tr>
    <tr>
      <td><a href="https://huggingface.co/bigcode/starcoder2-15b"><i class="fas fa-fw fa-solid fa-link"></i></a> StarCoder2-15B</td>
      <td>15B</td>
      <td>46.3</td>
      <td>61.6</td>
      <td>23.1</td>
    </tr>
    <tr>
      <td><a href="https://huggingface.co/bigcode/starcoderbase"><i class="fas fa-fw fa-solid fa-link"></i></a> StarCoderBase</td>
      <td>15.5B</td>
      <td>30.3</td>
      <td>51.7</td>
      <td>18.6</td>
    </tr>
    <tr>
      <td><a href="https://huggingface.co/codellama/CodeLlama-34b-hf"><i class="fas fa-fw fa-solid fa-link"></i></a> CodeLlama-34B</td>
      <td>34B</td>
      <td>45.1</td>
      <td>54.0</td>
      <td>10.2</td>
    </tr>
    <tr>
      <td><a href="https://huggingface.co/Phind/Phind-CodeLlama-34B-v2"><i class="fas fa-fw fa-solid fa-link"></i></a> Phind-V2</td>
      <td>34B</td>
      <td>71.9</td>
      <td>65.6</td>
      <td>32.1</td>
    </tr>
    <tr>
      <td><a href="https://ai.google.dev/models/gemini"><i class="fas fa-fw fa-solid fa-link"></i></a> Gemini-Pro</td>
      <td>—</td>
      <td>67.7</td>
      <td>59.3</td>
      <td>25.1</td>
    </tr>
    <tr>
      <td><a href="https://platform.openai.com/docs/models"><i class="fas fa-fw fa-solid fa-link"></i></a> GPT-3.5</td>
      <td>—</td>
      <td>61.5</td>
      <td>76.0</td>
      <td><strong>39.6</strong></td>
    </tr>
    <tr>
      <td><a href="https://platform.openai.com/docs/models"><i class="fas fa-fw fa-solid fa-link"></i></a> GPT-4</td>
      <td>—</td>
      <td><strong>84.1</strong></td>
      <td><strong>76.1</strong></td>
      <td>37.8</td>
    </tr>
  </tbody>
</table>

<p class="footnote">Last updated March 5, 2024</p>

<p>If you would like a model added you can reach out to <a href="mailto:dnicho@umd.edu">dnicho@umd.edu</a> or open an
<a href="https://github.com/parallelcodefoundry/ParEval/issues">issue in the GitHub repo</a>.</p>

<h2 id="citing-pareval">Citing ParEval</h2>

<div class="language-plaintext bibtex-block highlighter-rouge"><div class="highlight"><pre class="highlight"><code>@misc{nichols2024large,
      title={Can Large Language Models Write Parallel Code?}, 
      author={Daniel Nichols and Joshua H. Davis and Zhaojun Xie and 
              Arjun Rajaram and Abhinav Bhatele},
      year={2024},
      publisher = {Association for Computing Machinery},
      address = {New York, NY, USA},
      booktitle = {Proceedings of the 33rd International Symposium on High-Performance Parallel and Distributed Computing},
      series = {HPDC '24}
}
</code></pre></div></div>]]></content><author><name>UMD HPC Group</name></author><category term="LLMs for HPC" /><category term="LLM" /><category term="parallel" /><category term="benchmark" /><summary type="html"><![CDATA[We introduced the ParEval benchmark in “Can Large Language Models Write Parallel Code?” to evaluate the capability of LLMs at parallel code generation. We found a significant gap between their ability to generate sequential vs parallel code for a large array of computational problems and parallel programming models. On this page we keep an up-to-date table tracking the progress of state-of-the-art LLMs on ParEval.]]></summary></entry></feed>