We’re currently very busy writing down our projects, so you have a better idea what we can do.
As we were too busy handling all the projects and the required growth in the past 5+ years, we did not share much what we did.
By repeated request, we’re now writing these down, so you have an idea of what we are capable of. Please come back later for more stories, or follow us on LinkedIn.
The Fastest Payroll System Of The World
At StreamHPC we do several very different types of projects, but this project has been very, very different. In the first place, it was nowhere close to scientific simulation or media processing. Our client, Intersoft solutions, asked us to speed up thousands of payroll calculations on a GPU. They wanted to solve a simple problem, avoiding slow conversations with HR of large companies: Yes, I can answer your questions. For that I need to do a test-run. Please come back tomorrow. The calculation of 1600 payslips took one hour. This means 10,000 employees would take over 6 hours. Potential customers…
We accelerated the OpenCL backend of pyPaSWAS sequence aligner
Last year we accelerated the OpenCL-code in PaSWAS, which is open source software to do DNA/RNA/protein sequence alignment and trimming. It has users world-wide in universities, research groups and industry. Below you’ll find the benchmark results of our acceleration work. You can also test out yourself, as the code is public. In the readme-file you can learn more about the idea of the software. Lots of background information is described in these two papers: We chose PaSWAS because we really like bio-informatics and computational chemistry – the science is interesting, the problems are complex and the potential GPU-speedup is real.…
Learn about AMD’s PRNG library we developed: rocRAND – includes benchmarks
When CUDA kept having a dominance over OpenCL, AMD introduced HIP – a programming language that closely resembles CUDA. Now it doesn’t take months to port code to AMD hardware, but more and more CUDA-software converts to HIP without problems. The real large and complex code-bases only take a few weeks max, where we found that solved problems also made the CUDA-code run faster. The only problem is that CUDA-libraries need to have their HIP-equivalent to be able to port all CUDA-software. Here is where we come in. We helped AMD make a high-performance Pseudo Random Generator (PRNG) Library, called…
Demo: cartoonizer on an Altera Arria 10 FPGA
It takes quite some effort to program FPGAs using VHDL or Verilog. Since several years Intel/Altera has OpenCL-drivers, with the goal to reduce this effort. OpenCL-on-FPGAs reduced the required effort to a quarter of the time, while also making it easier to alter the specifications during the project. Exactly the latter was very beneficiary when creating the demo, as the to-be-solved problem was vaguely defined. The goal was to make a video look like a cartoon using image filters. We soon found out that “cartoonized” is a vague description, and it took several iterations to get the right balance between…
Bug fixing the MESA 3D drivers
Most of our projects are around performance optimisation, but we’re cleaning up bugs too. This is because you can only speed up software when certain types of bugs are cleared out. A few months ago, we got a different type of request. If we could solve bugs in MESA 3D that appear in games. Yes, we wanted to try that and got a list of bugs to solve. And as you can read, we were successful. Below you found a detailed description of one of the 5 bugs we solved by digging deep into the different games and the MESA…
Caffe and Torch7 ported to AMD GPUs, MXnet WIP
Last week AMD released ports of Caffe, Torch and (work-in-progress) MXnet, so these frameworks now work on AMD GPUs. With the Radeon MI6, MI8 MI25 (25 TFLOPS half precision) to be released soonish, it’s ofcourse simply needed to have software run on these high end GPUs. The ports have been announced in December. You see the MI25 is about 1.45x faster then the Titan XP. With the release of three frameworks, current GPUs can now be benchmarked and compared. Especially the expected good performance/price ratio will make this very interesting, especially on large installations. Another slide discussed which frameworks will…
We have been awarded the Khronos project to upgrade the OpenCL test suite to 2.2!
Some weeks ago we started with implementing the Compiler Test Suite for OpenCL 2.2. The biggest improvement of OpenCL 2.2 is C++ kernels, which originally was planned for 2.1. SPIRV 1.1 is another big improvement. We are very happy to have a part in making OpenCL better! We find OpenCL C++ kernels very important, even if it has its limitations. Thanks to SPIRV 1.1 it gets easier to have more (unofficial) kernel languages next to C and C++, and to get SYCL. Also upgrading from 2.0 to 2.2 is rather easy thanks to the open source libclcxx. Personally I found this project to…
How we sped up a flooding simulation 35 times (from 32-core CPU to multi-GPU)
How water moves through an area given a certain pace of instream, can be fully simulated. We got a request to make such simulation faster, as it took already too much time to do moderate simulations. As the customer wanted to be able to have more details, larger areas and more alternative situations computed, the current performance did not suffice. The code was already ported to MPI to scale to 8 cores. This code was used as a base for creating our optimised GPU-code. Using a single GPU we managed to get an 44 to 58 times speedup over single core…
Porting Manchester’s UNIFAC to OpenCL@XeonPhi: 160x speedup
As we cannot use the performance results for most of our commercial projects because they contain sensitive data, we were happy that Dr. David Topping from the University of Manchester was so kind to allow us to share the data for the UNIFAC project. The goal for this project was simple: port the UNIFAC algorithm to the Intel XeonPhi using OpenCL. We got a total of 485x speedup: 3.0x for going from single-core to multi-core CPU, 53.9x for implementing algorithmic, low-level improvements and a new memory layout design, and 3.0x for using the XeonPhi via OpenCL. To remain fair, we used the 160x speedup from…
We ported GROMACS from CUDA to OpenCL
GROMACS is an important molecular simulation kit, which can do all kinds of “soft matter” simulations like nanotubes, polymer chemistry, zeolites, adsorption studies, proteins, etc. It is being used by researches worldwide and is one of the bigger bio-informatics softwares around. To speed up the computations, GPUs can be used. The big problem is that only NVIDIA GPU could be used, as CUDA was used. To make it possible to use other accelerators, we ported it to OpenCL. It took several months with a small team to get to the alpha-release, and now I’m happy to present it to you. For…