Research direction 02
Malware Detection & Deep Learning
Malware detection that combines deep representation learning with heuristic search for feature selection, starting from careful replication of existing work.
- Status
- Baseline Investigation
- Lead
- Yoga
- Current stage
- Studying and replicating an external published approach.
Context
Malware detection systems work with high-dimensional, noisy features extracted from executables and their behaviour. Which features a model sees, and how they are represented, shapes what it can learn.
Deep representation learning can compress raw features into more informative representations. Heuristic search algorithms can explore the space of feature subsets that would be impractical to search exhaustively. Hybrid approaches combine the two.
Before proposing anything new, this direction starts by understanding and replicating an existing hybrid approach. A reliable baseline is a precondition for any later claim of improvement.
Research areas
- Malware Detection
- Classifying software samples as malicious or benign from extracted static or behavioural features.
- Deep Representation Learning
- Learning compact representations of input features with neural networks.
- Feature Selection
- Choosing a subset of features that preserves useful signal while reducing dimensionality.
- Heuristic Search
- Search strategies that explore large combinatorial spaces, such as feature subsets, without exhaustive enumeration.
- Reproducible Experiments
- Fixed seeds, documented configurations, and recorded deviations, so results can be checked and repeated.
Current focus
Baseline replication
The research foundation is an external paper: “Improving malware detection performance using hybrid deep representation learning with heuristic search algorithms.” It is under investigation for baseline replication. It was not authored by Yoga or any AsteriaX Labs collaborator.
The aim at this stage is to reproduce the described approach as faithfully as the publication allows, and to document every point where the description leaves implementation choices open.
Thesis novelty has not been defined. That decision depends on what the replication shows.
Fig. 02 · Hybrid detection pipeline
Questions guiding the replication
- Q1
Can the described pipeline be reproduced from the published information alone?
- Q2
Which implementation details are underspecified, and how much do reasonable choices for them matter?
- Q3
How stable are selected feature subsets across random seeds and data splits?
- Q4
What does a well-defined baseline need to include so that later comparisons are meaningful?
Replication plan
A plan of work, not a record of completed steps.
Study the publication
Extract the pipeline, data handling, and evaluation protocol, and list every unstated assumption.
Prepare data
Assemble the dataset and preprocessing as closely to the paper’s description as possible.
Implement the representation stage
Build the deep representation learning component and verify its behaviour in isolation.
Implement heuristic feature selection
Implement the search procedure over feature subsets and record its configuration.
Evaluate and document
Run the documented protocol across seeds and record agreements and discrepancies with the publication.
Scope and claims
What this work investigates
- Whether an existing hybrid approach can be reproduced reliably.
- Sensitivity of the approach to implementation choices the publication does not fix.
- What a sound, reusable baseline for later work looks like.
What it does not claim
- Authorship of the paper under investigation.
- A defined thesis contribution or novelty. This is not yet decided.
- Benchmark results or completed experiments.