Graph-Based Deep Learning

Authors

  • Ivan Silva Author
  • Laura Kovacs Author

DOI:

https://doi.org/10.5281/zenodo.19614710

Keywords:

graph neural networks; message passing; graph transformers; molecular graphs; node classification; link prediction; equivariant networks; over-squashing

Abstract

Graph-structured data pervades science and industry -- social networks, molecular graphs, knowledge bases, communication networks, and biological interaction maps all naturally represent entities as nodes and relationships as edges -- yet standard deep learning architectures (CNNs, transformers) assume grid or sequence structure and cannot directly process irregular graph topologies. Graph neural networks address this by learning node representations through iterative message passing between connected nodes, capturing both node features and graph topology in a unified framework. This study presents a controlled comparison of six GNN architectures -- GCN, GAT, GraphSAGE, GIN, GPS (graph transformer), and equivariant GNNs (EGNN) -- across four graph learning tasks: node classification (citation networks: Cora, CiteSeer, PubMed), graph classification (molecular property prediction: OGBG-MolHIV, OGBG-MolPCBA), link prediction (knowledge graph completion: OGB-WikiKG2), and 3D molecular generation (QM9 property regression). A total of 2,040 experiments were conducted under standardised training protocols. GPS graph transformer achieved the highest performance on graph classification (AUROC = 0.812 +- 0.008 on MolHIV) by combining local message passing with global self-attention, capturing both neighbourhood structure and long-range graph dependencies. EGNN achieved the best molecular property prediction (MAE = 0.012 eV on QM9 HOMO-LUMO gap) through equivariance to 3D rotations and translations. On node classification, GAT (86.4 +- 0.4% on Cora) marginally outperformed GCN (85.8 +- 0.5%) through learned attention weights, but the gap was smaller than commonly reported due to careful hyperparameter tuning of baselines. The over-squashing problem -- information loss in deep GNNs caused by exponential neighbourhood growth -- limited all message-passing GNNs to 4-6 effective layers, with GPS overcoming this through global attention. A practical GNN architecture selection guide mapping graph type, task, and scale to recommended configurations is proposed.

Downloads

Published

2026-08-19

How to Cite