Writing

GCN_for_EEG, five years later

In short: GCN_for_EEG got its stars because it removed one barrier, MATLAB, at a time when graph networks on EEG were new to most people. Five years on, the idea still holds, but the setup would not…

published
read time
5 min
words
925
lang
en
filed under
Research

In short: GCN_for_EEG got its stars because it removed one barrier, MATLAB, at a time when graph networks on EEG were new to most people. Five years on, the idea still holds, but the setup would not survive a fresh install, and I would change how the graph is built and how results are reported.

In 2020 I was in the last year of a biomedical engineering degree, reading everything I could find about graph neural networks. I wrote up the ones I found most useful in a review of GNN architectures, and around the same time I put two small repositories on GitHub. They are still the most-starred things I have made.

85
stars on GCN_for_EEG
10
stars on MNE_GCN, the follow-up
64
EEG electrodes, one graph node each
4
motor imagery classes

What the repo actually is

It classifies four classes of motor imagery EEG from the PhysioNet EEG Motor Movement/Imagery dataset. Each of the 64 electrodes is a node. The edges come from how strongly channels move together, an absolute Pearson correlation matrix. A graph convolutional network then learns over that graph.

I did not invent any of the hard parts, and the README says so. The graph convolution comes from Michaël Defferrard's cnn_graph. The EEG pipeline comes from Shuyue Jia's EEG-DL, which is excellent, but its preprocessing ran in MATLAB. My contribution was a Python version of that preprocessing, a few extra GCN variants, and some changes to the code. MNE_GCN came later and swapped the loading step for MNE, so you could pull in other EEG datasets and change the band-pass filter yourself.

Why I think it got stars

I never ran a survey, so this is a guess. The README title says it: the pure Python interpretation. A lot of people, researchers and newcomers alike, who wanted to try GCNs on EEG did not have a MATLAB licence, and this was a way in without one. GCNs on brain signals were also a fresh idea to many people in 2020, and a repo that runs end to end on a public dataset is worth more to a newcomer than a paper.

The stars were not for the model. They were for removing a barrier.

Where people get stuck

I do not keep a tally, but the questions I remember were about getting it to run, not about graphs. Reading the README today, I can see why. Getting from download to result takes four manual stages, two Python versions, and copying files between folders by hand.

2020 download edfread.py Python 2.7 preprocess 128 .mat files onEEGcode.py TF 1.13 copy by hand copy by hand today one script, one config, one environment
Every hand-off between boxes is a place where a new user gives up.

The dependency list is the other problem. TensorFlow 1.13 and Python 2.7 were already old in 2020. Today they are hard to install at all on a new machine.

What I would rewrite

PartIn 2020Now
EnvironmentPython 2.7 for reading, TensorFlow 1.13 for trainingPython 3 and PyTorch Geometric, pinned in one file
Loadingedfread.py, then copy files between foldersMNE reads the EDF files directly, as MNE_GCN started to do
PipelineNumbered folders, run in orderOne command with a config file
GraphAbsolute Pearson matrixSame idea, built from training trials only
SplitNot obvious from the READMEBy subject, stated up front
ResultsAn accuracy curvePer-subject scores, a confusion matrix, and a CSP baseline

Two rows deserve a sentence each.

The graph. A correlation matrix is computed from data. If that data includes the test trials, the graph has already seen the test set, a quiet form of leakage. I would want to check where the matrix comes from before trusting any number built on it, and in a rewrite I would build it from training trials only:

import numpy as np

def eeg_graph(train_trials, k=8):
    """train_trials: (n_trials, n_channels, n_samples), training data only."""
    n_ch = train_trials.shape[1]
    x = train_trials.transpose(1, 0, 2).reshape(n_ch, -1)
    a = np.abs(np.corrcoef(x))           # absolute Pearson, channel by channel
    np.fill_diagonal(a, 0)

    # keep each channel's k strongest neighbours, then make it symmetric
    keep = np.argsort(a, axis=1)[:, -k:]
    mask = np.zeros_like(a, dtype=bool)
    np.put_along_axis(mask, keep, True, axis=1)
    a = np.where(mask | mask.T, a, 0.0)

    # the usual GCN normalization: D^-1/2 (A + I) D^-1/2
    a_hat = a + np.eye(n_ch)
    d = 1.0 / np.sqrt(a_hat.sum(axis=1))
    return a_hat * d[:, None] * d[None, :]

The normalization at the end is the one from the review post. Without it, channels with many strong neighbours dominate for no good reason.

The baseline. I had already written Common Spatial Patterns for motor imagery in CSP_on_EEG that same year. A rewrite would report CSP with a linear classifier next to the GCN, on the same subject-wise split. If the graph network does not beat that, the graph is decoration.

What I would keep

The core idea: electrodes as nodes, relationships between channels as edges. It matches how EEG works, where what one channel sees is shaped by its neighbours. I would also keep the honesty of the README, which credits the people whose work it stands on in the first lines. That is the part I am proudest of, looking back.

If you are using it today

Do not fight Python 2.7. Start from MNE_GCN for loading, port the model to PyTorch Geometric, which has the graph convolution built in, and build your adjacency matrix from training subjects only. Then run a CSP baseline on the same split before you look at the GCN's score. If you get it working on a modern stack, open a pull request. I would happily link to it from the README.

related

Keep reading