Projects

GPU Neighbor Gather: Divergence vs. Coalescing

Year
2026
Status
complete
Built with
CUDA, C++

Technical summary

A benchmark of the neighbour-gather pattern at the heart of particle simulations and graph neural networks. It varies how neighbour counts are distributed across elements and measures whether warp divergence or the memory-access pattern dominates the cost, a question raised by OpenFPM's claims of near hand-tuned SPH throughput.