Research direction 01
Federated Learning & Cybersecurity
How federated intrusion detection systems share knowledge across clients whose traffic differs, with a focus on prototype-based methods and rare attack classes.
- Status
- Ongoing Research
- Lead
- Will
- Current stage
- Thesis research in progress within an existing published framework.
Context
Network intrusion detection benefits from data held by many independent parties, but that data is sensitive and rarely leaves its origin. Federated learning lets participants train collaboratively while keeping raw traffic local.
In practice, client data is non-IID. Each participant observes a different mix of benign traffic and attack classes, and some classes appear at only a few sites. These conditions make aggregation harder and tend to disadvantage rare classes.
Prototype-based knowledge sharing is one response. Instead of exchanging full model parameters, clients share compact class-level representations, called prototypes, which a server aggregates and redistributes.
Research areas
- Federated Intrusion Detection
- Training intrusion detection models across multiple network operators or sites without centralising raw traffic.
- Non-IID Data
- Client datasets that differ in class balance and feature distribution, which is the normal case for real networks.
- Prototype-Based Knowledge Sharing
- Exchanging per-class representative embeddings between clients and server rather than full model weights.
- Server-Side Prototype Aggregation
- How the server combines prototypes from different clients into a global prototype for each class.
- Rare-Class Detection
- Detecting attack classes that are under-represented overall or present at only a small number of clients.
Current focus
Class-support-weighted prototype aggregation in PROTEAN
The current thesis research works within PROTEAN, an existing published method for prototype-based federated intrusion detection. PROTEAN is the subject of study and the baseline. It is not an AsteriaX Labs invention.
The investigation modifies how the server weights client prototypes during aggregation. In the formulation under study, a client’s contribution to a class prototype is weighted by its support for that class, raised to a tunable exponent β. The rest of the framework is retained, so that observed differences can be attributed to the aggregation change alone.
β sets how strongly support matters. At β = 0, every contributing client is weighted equally. As β increases, clients holding more samples of a class have more influence on that class’s global prototype. Whether, and when, this helps is the open question.
Fig. 01 · Server-side prototype aggregation
Research questions
- Q1
How does weighting client prototypes by class support change the aggregated prototypes under different degrees of non-IID partitioning?
- Q2
How sensitive are detection outcomes, especially for rare classes, to the choice of β?
- Q3
Under which data distributions does support weighting help, make no measurable difference, or hurt?
- Q4
What evaluation protocol makes comparison against unmodified PROTEAN fair and reproducible?
Method outline
A plan of work, not a record of completed steps.
Reproduce the baseline
Run PROTEAN as published to establish a reference point under a fixed, documented configuration.
Define partitioning scenarios
Construct client splits with controlled levels of class imbalance and class absence.
Introduce support weighting
Change only the server-side aggregation weights, leaving local training and communication untouched.
Sweep β
Evaluate a range of exponents, including β = 0, across every scenario and multiple seeds.
Report per class
Analyse results per class rather than only in aggregate, so effects on rare classes are visible.
Scope and claims
What this work investigates
- A specific, contained modification to server-side prototype aggregation.
- The effect of a tunable exponent β on class-support weighting.
- Behaviour under controlled non-IID partitions, with attention to rare classes.
What it does not claim
- That support-weighted aggregation universally improves detection performance.
- Authorship of PROTEAN, which is an existing published method.
- Any results. Findings will be reported only once experiments are complete and reviewed.