Where We Share Many Things We Learn on the Job

Programming and maxing out GPUs

We make software that runs really fast on GPUs, mostly focused on compute, mathematics and visuals. Most blogs are therefore about software performance engineering.

Handy tools and tricks

We need quality software to get performant software. A lot of tricks and tools are useful outside our domain, and we share them here.

Computer hardware and compilers

How Color Really Works: An Introduction to Color Spaces

When working with images as a software engineer, color may seem a straightforward concept. After decoding an image you typically end up with an array of pixel values, each a tuple of three integer components. These components generally are “RGB” and so the value (255, 0, 0) must denote red,…

Keep reading…

HIP Compiler Acceleration

It’s one of the age-old problems of software development: You’ve started a new project, set up fancy CI infrastructure, and everything works well. But as time goes on, you add more and more features to the project. Compile-times on your local checkout of the projects are mostly fine as you…

Keep reading…

When to Blame the Compiler: An NVCC Case Study

In our line of work, performing low-level performance optimization, the compiler is less of a language abstraction and more of a tool to directly interact with. Compilers are complex machines. Correctly converting C++-like code to assembly instructions is a gigantic challenge, producing optimized instructions even more so. It is therefore…

Keep reading…

RDNA and CDNA: Similarities and Differences

In 2019 AMD announced the Radeon™ RX 5700 XT, a GPU that sported its brand-new at the time architecture named RDNA. It aimed to provide upgrades compared to the older GCN-based cards. Then one year later, AMD announced another GPU architecture — CDNA, with the release of the Radeon™ Instinct…

Keep reading…

Asynchronous and Parallel Programming in C++26

For the past decade, every C++ programmer who wanted to do real concurrent work has had the same conversation with themselves. std::async is a toy. std::thread is too low-level. std::future doesn’t compose. So you reach for TBB, or another third-party library, or a thread pool you wrote in 2017, or…

Keep reading…

Lazy Ranges in C++23 with Std::generator

Disclaimer: LLMs were used for proof-reading and grammar check. C++20 gave us coroutines. The machinery needed was included: co_yield, co_return, co_await but the standard library had no concrete coroutine types. You had to write your own promise type, your own iterator, your own bookkeeping. The complexity of this boilerplate was…

Keep reading…

Our Two Presentations at GPU Day 2026

At Stream HPC, we enjoy opportunities to connect with the HPC and accelerator community, exchange ideas, and learn from engineers and researchers working across the GPU ecosystem. Later this month, several members of our team will be attending GPU Day 2026 in Budapest, Hungary. Now in its 16th edition, GPU…

Keep reading…

Thinking with iterators in CUDA and HIP

Parallel primitives are the ubiquitous building blocks of GPU programming with CUDA and HIP, to make your life as a programmer easier. Primitives like scans, reductions, and sorts operate in parallel over large data inputs. The basic use case has input and output residing in device memory as an array…

Keep reading…

Loading…

Something went wrong. Please refresh the page and/or try again.

Ready to Optimize?

Enhance Your Software Performance with Our Experts

Contact us today to discuss how Stream HPC can elevate your computing capabilities and lower costs.