The future Large Hadron Collider (LHC) will produce much more data than today. A new method using artificial intelligence could help physicists sort through it.
The LHC collides protons to study the smallest components of matter. Its future version is scheduled to go into operation in 2030 at the earliest. It will produce about 40 million collisions per second, with four to five times more simultaneous collisions than the current maximum.

The CMS experiment at the LHC
Image: CERN
The CMS detector will have to find the particles produced amid this activity. It will receive HGCAL, a device equipped with about six million sensors. When a particle passes through it, it creates a cascade of new particles. This group is called a shower.
Each sensor records a small part of these showers and their energy. The problem arises when several cascades mix together. The software must then link each signal to the correct original particle. This task becomes especially difficult when collisions occur almost at the same location.
Current methods try to reconstruct each shower directly from its characteristics. The new method using AI proceeds differently. It first learns which signals resemble one another and could share the same origin. It simultaneously separates those that seem to come from different particles.
A second step then groups the associated signals. This technique is called contrastive learning, because it learns by comparing similarities and differences. The researchers compared this method with the current system on fully simulated data.
In these tests, the new approach better separated overlapping showers. It also estimated their energy more accurately. Its advantage grew in the busiest events. It even remained effective with amounts of particles and energies that had not been part of its training.
The algorithm must now succeed in a complete simulation of CMS and HGCAL. Additional tests will then determine whether this method can actually analyze future LHC collisions.