AMD RESEARCH & ADVANCED DEVELOPMENT · AUSTIN, TEXAS

Vignesh
Adhinarayanan

Architecture and systems for large-scale AI inference.

I research GPU architecture, high-bandwidth memory, interconnection networks, data orchestration, and parallel execution models for AI accelerators.

GPGPUsDistributed inferenceHBMCollectives
20+peer-reviewed publications
13U.S. patent grants
3international patent grants
ISCA publications

RESEARCH AGENDA

Co-design across the full inference stack.

My work connects model structure to the physical costs of computation, data movement, memory capacity, and synchronization.

01

Interactive AI inference

Performance models and execution strategies for frontier language models, speculative decoding, long-context inference, and mixture-of-experts architectures.

MODEL × SYSTEM
02

GPU communication

Low-latency collectives, on-chip and scale-out interconnects, and topology-aware mappings for distributed accelerator systems.

COMPUTE × NETWORK
03

Memory systems

High-bandwidth memory organizations and near-memory techniques for fine-grained, irregular, and bandwidth-constrained workloads.

ARCHITECTURE × CIRCUITS

SELECTED WORK

Recent publications

Research spanning AI inference efficiency, 3D-stacked memory, and large-scale computing systems.

Complete publication record

CURRENT FOCUS

Maximizing interactivity for frontier models through hardware and software co-design.

Contact →