TECH BLOG
Where We Share Many Things We Learn on the Job
Programming and maxing out GPUs
We make software that runs really fast on GPUs, mostly focused on compute, mathematics and visuals. Most blogs are therefore about software performance engineering.
Handy tools and tricks
We need quality software to get performant software. A lot of tricks and tools are useful outside our domain, and we share them here.
Computer hardware and compilers
To make the best software, often 2 protagonists are in the way: the hardware and the compiler.
OpenGL “Compute Shaders” Before Compute Shaders – Heroic Times
The times before the advent of compute shaders, twenty years ago, were heroic. GPU came out of infancy into adolescence: the pace of innovation was breathtaking. A year without new GPU architecture was unthinkable. You could see new programming languages and APIs appearing one after another, bringing us closer and…
Keep reading…
How Color Really Works: An Introduction to Color Spaces
When working with images as a software engineer, color may seem a straightforward concept. After decoding an image you typically end up with an array of pixel values, each a tuple of three integer components. These components generally are “RGB” and so the value (255, 0, 0) must denote red,…
Keep reading…
HIP Compiler Acceleration
It’s one of the age-old problems of software development: You’ve started a new project, set up fancy CI infrastructure, and everything works well. But as time goes on, you add more and more features to the project. Compile-times on your local checkout of the projects are mostly fine as you…
Keep reading…
When to Blame the Compiler: An NVCC Case Study
In our line of work, performing low-level performance optimization, the compiler is less of a language abstraction and more of a tool to directly interact with. Compilers are complex machines. Correctly converting C++-like code to assembly instructions is a gigantic challenge, producing optimized instructions even more so. It is therefore…
Keep reading…
The LuaJIT NYI That Silently Poisoned an Unrelated Hot Loop
If you haven’t used Lua before, it’s the go-to embedded scripting language for games like Factorio and World of Warcraft, and for applications like Neovim and OpenResty. LuaJIT, its just-in-time compiler, is widely praised for its speed, so it’s easy to assume your code is already running as fast as…
Keep reading…
Image Upscaling Using Convolutional Neural Networks in Vulkan
In the world of high-performance programming, true portable code has been a goal for a long time. Vulkan has been very successful at this, with Vulkan drivers being provided by every major GPU vendor and available for all common desktop and mobile operating systems. Platforms without first party support (e.g.…
Keep reading…
Kernel Development + C++26 Reflection = 🧡
Anyone who has written a non-trivial GPU kernel has hit the same wall. You define a struct in C++ and pass it to a kernel. Six months later, somebody adds a member, forgets to update the matching layout on the device side, and a 200-line shader reads garbage. The compiler…
Keep reading…
RDNA and CDNA: Similarities and Differences
In 2019 AMD announced the Radeon™ RX 5700 XT, a GPU that sported its brand-new at the time architecture named RDNA. It aimed to provide upgrades compared to the older GCN-based cards. Then one year later, AMD announced another GPU architecture — CDNA, with the release of the Radeon™ Instinct…
Keep reading…
Asynchronous and Parallel Programming in C++26
For the past decade, every C++ programmer who wanted to do real concurrent work has had the same conversation with themselves. std::async is a toy. std::thread is too low-level. std::future doesn’t compose. So you reach for TBB, or another third-party library, or a thread pool you wrote in 2017, or…
Keep reading…
Lazy Ranges in C++23 with Std::generator
Disclaimer: LLMs were used for proof-reading and grammar check. C++20 gave us coroutines. The machinery needed was included: co_yield, co_return, co_await but the standard library had no concrete coroutine types. You had to write your own promise type, your own iterator, your own bookkeeping. The complexity of this boilerplate was…
Keep reading…
Our Two Presentations at GPU Day 2026
At Stream HPC, we enjoy opportunities to connect with the HPC and accelerator community, exchange ideas, and learn from engineers and researchers working across the GPU ecosystem. Later this month, several members of our team will be attending GPU Day 2026 in Budapest, Hungary. Now in its 16th edition, GPU…
Keep reading…
Thinking with iterators in CUDA and HIP
Parallel primitives are the ubiquitous building blocks of GPU programming with CUDA and HIP, to make your life as a programmer easier. Primitives like scans, reductions, and sorts operate in parallel over large data inputs. The basic use case has input and output residing in device memory as an array…
Keep reading…Loading…
Something went wrong. Please refresh the page and/or try again.
Enhance Your Software Performance with Our Experts
Contact us today to discuss how Stream HPC can elevate your computing capabilities and lower costs.