Teledyne SP Devices has demonstrated convolutional neural-network inference directly on a high-speed digitiser operating at 10 giga-samples per second, processing a 120Gbit/s detector stream with an inference latency of 300 nanoseconds. Developed through doctoral research at Fulda University of Applied Sciences with GSI Helmholtzzentrum für Schwerionenforschung, the system places the neural network on the ADQ35 digitiser’s onboard AMD Kintex UltraScale KU115 FPGA rather than transferring the complete signal elsewhere for analysis.
The research targets particle pile-up, where several detector pulses arrive close enough together for their electrical signals to overlap. Conventional peak detection can then struggle to determine how many particles generated the combined waveform and when each event occurred, particularly as beam intensity rises. Reducing the particle rate makes separation easier but also limits the amount of experimental data collected, creating an incentive to interpret the denser signal rather than avoid it.
The neural network has been trained to identify the individual events represented inside those overlapping waveforms and extract particle count and time-of-arrival information. Because inference occurs directly behind the analogue-to-digital conversion, the system does not need to move the full-rate raw stream through an external computing layer before producing a result, avoiding a data-transfer stage that would itself require exceptional bandwidth and add latency.
At 10GSPS and 12-bit resolution, the acquisition path represents 120Gbit/s of information before wider system overheads are considered. Teledyne says the FPGA implementation handles 32 samples in parallel at 312.5MHz, allowing the processing pipeline to sustain the complete sampling rate while returning an inference after roughly 300ns. That deterministic response is possible because the model is mapped into hardware logic rather than scheduled as a general-purpose software workload.
FPGA implementation changes the balance between speed and flexibility. A GPU can generally run different models through software with less hardware redesign, whereas an FPGA network has to be synthesised for the target architecture; once implemented, however, operations can be arranged as parallel pipelines with predictable timing. That makes the device useful where the answer has to be available within the measurement cycle rather than after data has been stored and analysed offline.
Resource consumption determines whether the AI can coexist with the rest of the instrument. Teledyne reports that the CNN uses less than 15% of the available digital signal-processing resources and less than 6% of the FPGA’s lookup tables, leaving capacity for acquisition, buffering, triggering or additional algorithms. A model that consumed most of the device could still prove the concept, but integrating it into a practical measurement platform would become considerably harder.
Hardware-aware training and implementation are therefore central to the result. The network has to deliver sufficient detection performance while fitting the parallelism, memory and arithmetic available on the FPGA, rather than beginning with an unconstrained AI model and discovering later that the hardware cannot execute it at the required rate. That design discipline is one reason the research has relevance beyond the particle-physics application itself.
High-speed RF measurement, scientific instrumentation and some industrial inspection systems face a similar problem when sensors generate data more quickly than it is convenient to transfer everything to a central computer. Performing inference close to the acquisition point can reduce the amount of information that needs to move through the rest of the system, provided the local model can make the required decision accurately and within a known latency.
The GSI work remains research rather than a production installation sold as an AI system. Teledyne has also been explicit that UFAIRA, a 2026 spinout based on related FPGA AI methods, was not a party to the original research and that GSI is not described as a UFAIRA customer. Keeping those distinctions intact matters because a successful demonstrator shows technical feasibility without establishing how widely the approach has been commercialised.
The measured result is nevertheless substantial: a commercial digitiser is acquiring and analysing 120Gbit/s in real time, separating overlapping detector events with a CNN and producing an answer within 300ns while using a minority of the available FPGA resources. Replicating that architecture in other applications will require different training data and models, but the experiment shows that AI inference can be moved into the acquisition hardware without sacrificing the deterministic timing on which many high-speed instruments depend.



