Lune

ISCA2025顶会

Nyx: Virtualizing dataflow execution on shared FPGA platforms

Panagiotis Miliadis, Dimitris Theodoropoulos, Nectarios Koziris, Dionisios N. Pnevmatikatos

2025年份
1顶会引用

摘要

As FPGAs become more widespread for improving computing performance within cloud infrastructure, researchers aim to equip them with virtualization features to enable resource sharing in both temporal and spatial domains, thereby improving hardware utilization.Existing multi-tenant solutions focus on task-parallel models, where tasks are assigned to distinct regions to process separate sets of data.However, this model introduces waiting times between dependent and pipelined tasks, leading to longer response times for applications.The root cause is the lack of support for dataflow execution -a key potential of FPGAs and a crucial optimization for applications.Dataflow allows direct data streaming between operators, forming a task-pipelined model that reduces application latency by overlapping task operations within its workflow.This paper presents Nyx, the first system to enable dataflow execution in a task-based virtualized and shared FPGA environment.Nyx enables efficient resource sharing by dividing the FPGA into distinct reconfigurable regions.At its core, Nyx employs virtual FIFOs, independent channels that allow seamless communication between pipelined tasks.Its approach ensures smooth task operation even when the predecessor or successor tasks are not simultaneously scheduled in the FPGA, making them agnostic to their dependencies, communication channels or data locations.An FPGA hypervisor is designed to handle all data dependencies and efficiently dispatch pipelined tasks across regions at high throughput.Nyx outperforms existing state of the art virtualized task-parallel approaches by 1.26x -8.87x across a series of real-world benchmarks.Furthermore, it reduces response times by 2.8x -3.28x during low-demand periods, decreasing also deadline violations by up to 76.5%.Under highdemand conditions, Nyx delivers 2x -2.75x reduction, 34.5% fewer violations, and up to 1.9x reduced tail response time.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get f5ac59bf-03d0-4359-9ebe-9dfc677944aa

引用它的顶会 Paper1

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖