Nevertheless, the researchers additionally examined Nvidia server GPUs resembling A100 and H100 or newer GPUs on the Blackwell structure like RTX 5090 or RTX 6000, and their assault didn’t produce bit flips. It is because these playing cards use totally different or newer sort of reminiscence resembling HBM, GDDR6X, and GDDR7, which have totally different defenses. The analysis workforce plans to research these chips sooner or later in order that they don’t low cost the chance that various assault patterns might exist for them.
Why does this matter?
Within the age of AI fashions, enterprise and server-class GPUs are worthwhile as a result of they’re wanted for each coaching or fine-tuning AI fashions and for working them, often known as inference. Even with out AI, these GPUs are normally put in in datacenters and run delicate workloads usually from a number of digital machines on the identical time.
Of their exams, the researchers managed to crash GPUs so usually that inside someday the playing cards flagged themselves as faulty and due for alternative utilizing their inner crash detection mechanisms. Apart from triggering denial-of-service circumstances that kill all of the workloads working on the cardboard, the researchers managed to deprave the GPU’s reminiscence web page tables in a means that allowed an unprivileged program to escalate its privileges to root.


