Portrait of Nic McDonald

Nic McDonald

Software and Hardware Engineer

NVIDIA

Biography

I am a computer architecture researcher and software/hardware engineer specializing in AI systems, accelerator design, and high-performance interconnects.

As a Senior Research Scientist at NVIDIA, I develop analytical models and simulation tools to evaluate LLM training and inference across compute, memory, and communication. My work spans hardware/software co-design, energy-efficient inference architectures, and electrical and optical networks. Previously, at Google and Hewlett Packard Labs, I developed network architectures, routing algorithms, transport protocols, and congestion-control systems.

I build tools that connect workload behavior to architectural decisions, including Calculon for LLM systems and SuperSim for interconnection networks.

Interests

  • Computer Architecture
  • Software Engineering
  • Interconnection Networks
  • High Performance Computing
  • Simulation and Modeling
  • Embedded Systems
  • Artificial Intelligence
  • Algorithmic Trading

Education

  • Ph.D. in Electrical Engineering, 2016

    Stanford University

  • M.S. in Electrical Engineering, 2012

    University of Utah

  • B.S. in Computer Engineering, 2010

    University of Utah

Projects

Publications

Cite

@inproceedings{isaev2023calculon,
  title={{Calculon: a Methodology and Tool for High-Level Codesign of Systems and Large Language Models}},
  author={Isaev, Mikhail and McDonald, Nic and Dennison, Larry and Vuduc, Richard},
  booktitle={Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC)},
  year={2023},
  organization={ACM}
}

Cite

@inproceedings{isaev2023scaling,
  title={{Scaling Infrastructure to Support Multi-Trillion Parameter LLM Training}},
  author={Isaev, Mikhail and McDonald, Nic and Vuduc, Richard},
  booktitle={Proceedings of the Workshop on Architecture and System Support for Transformer Models (ASSYST)},
  year={2023}
}

Cite

@inproceedings{isaev2022paragraph,
  title={{ParaGraph: An application-simulator interface and toolkit for hardware-software co-design}},
  author={Isaev, Mikhail and McDonald, Nic and Young, Jeffrey and Vuduc, Richard},
  booktitle={51st International Conference on Parallel Processing (ICPP)},
  year={2022},
  organization={ACM}
}

Cite

@inproceedings{mcdonald2019hxrouting,
  title={{Practical and Efficient Incremental Adaptive Routing for HyperX Networks}},
  author={McDonald, Nic and Isaev, Mikhail and Flores, Adriana and Davis, Al and Kim, John},
  booktitle={Proceedings of the Conference on High Performance Computing Network, Storage and Analysis (SC)},
  year={2019},
  organization={ACM}
}

Cite

@inproceedings{domke2019hyperx,
  title={{HyperX Topology: First At-Scale Implementation and Comparison to the Fat-Tree}},
  author={Domke, Jens and Matsuoka, Satoshi and Ivanov, Ivan R. and Tsushima, Yuki and Yuki, Tomoya and Nomura, Akihiro and Miura, Shin'ichi and McDonald, Nic and Floyd, Dennis L. and Dubé, Nicolas},
  booktitle={Proceedings of the Conference on High Performance Computing Network, Storage and Analysis (SC)},
  year={2019},
  organization={ACM}
}

Cite

@inproceedings{mcdonald2018supersim,
  title={{SuperSim: Extensible Flit-Level Simulation of Large-Scale Interconnection Networks}},
  author={McDonald, Nic and Flores, Adriana and Davis, Al and Isaev, Mikhail and Kim, John and Gibson, Doug},
  booktitle={International Symposium on Performance Analysis of Systems and Software (ISPASS)},
  year={2018},
  organization={IEEE}
}

Cite

@phdthesis{mcdonald2016hpsoc,
  title={{High-Performance Service-Oriented Computing}},
  author={McDonald, Nicholas},
  year={2016},
  school={Stanford University}
}

Experience

  • Feb 2021 – Present

    Senior Research Scientist

    NVIDIA
    Salt Lake City, Utah

    I work on AI system architecture at NVIDIA Research, using analytical models to explore software and hardware tradeoffs for large language models (LLMs). I developed Calculon, a framework for rapidly estimating performance and resource usage across a broad range of execution configurations, and built tools for large-scale LLM training analysis, including studies of models with up to 128 trillion parameters. This work was published at SC23 and helped identify more efficient LLM optimization and system design strategies. My research also examines scale-up and scale-out interconnects and network topologies, including approaches that use optical circuit switches for efficient large-scale training. I have contributed to distributed GEMM research and software infrastructure by defining reusable multi-GPU APIs and helping integrate distributed algorithms into CUTLASS. I also develop AI accelerator design concepts that minimize energy use and data movement while maximizing inference speed across emerging model workloads.

  • Jan 2019 – Jan 2021

    Senior Software Engineer

    Google
    Sunnyvale, California

    At Google, I developed next-generation data center network topologies, routing algorithms, transport protocols, congestion control systems, and switching architectures. This work included guiding the architecture of future accelerated networking hardware and software systems. I also developed ParaGraph, an application-simulator interface and toolkit for hardware/software co-design that connected compiler-derived workload models with simulation infrastructure to support large-scale system exploration. ParaGraph was published at ICPP 2022.

  • Jan 2016 – Dec 2018

    Research Scientist

    Hewlett Packard Labs
    Fort Collins, Colorado

    At Hewlett Packard Labs, I was a lead architect for a new high-performance network supporting large-scale high-performance computing (HPC) and massively parallel memory-driven computing (MDC) systems. I was a key designer of the Gen-Z routing architecture specification. Using data-driven simulation, I also guided the design of a novel multi-chip module (MCM) switch architecture that used co-packaged integrated photonics.

  • Aug 2008 – Sep 2012

    Digital Hardware Design Engineer

    L3 Communications
    Salt Lake City, Utah

    At L3 Communications, I developed digital processing architectures for encryption, networking, and waveform data processing. I designed systems to meet strict requirements for high performance, low power consumption, and a small hardware footprint. Each design underwent rigorous review for reliability and signal integrity.

Patents

Search