MARATTO

article · BMC Bioinformatics

A novel and innovative cancer classification framework through a consecutive utilization of hybrid feature selection

202339 citationsOpen accessDebre Tabor University

In plain language

Early cancer detection relies heavily on analysing complex gene expression datasets, which often feature thousands of genes across a relatively small number of patient samples. Identifying the most relevant genes is essential to speed up computational analysis and improve diagnostic accuracy. A hybrid computational framework combines two metaheuristic approaches, spider monkey optimisation and the cuckoo search algorithm, alongside a minimum redundancy maximum relevance data-cleaning technique to filter out redundant gene features. Deep learning models then classify these refined gene subsets into specific cancer types. Evaluated across eight benchmark microarray gene expression datasets, the framework was assessed using precision, recall, F1-score, and confusion matrices. The combined feature selection and deep learning pipeline achieved higher classification accuracy than existing machine learning and deep learning models across all tested datasets.

Key takeaways

  • A hybrid feature selection technique pairs spider monkey optimisation with the cuckoo search algorithm to isolate informative genes.
  • Data cleaning using minimum redundancy maximum relevance filters out redundant gene expression features to boost classification speed and precision.
  • Deep learning models classify the refined gene subsets to distinguish specific cancer types.
  • Testing on eight benchmark cancer microarray datasets demonstrated superior accuracy compared to conventional machine learning and deep learning methods.

Why it matters

Accurate early-stage cancer classification allows clinicians to select targeted treatments and improve patient outcomes. Gene expression profiles provide vital clues, but their high complexity often overwhelms conventional computing tools. By streamlining genetic data to only the most informative markers, computational methods can make molecular cancer diagnostics faster, more accurate, and more clinically useful.

Commercialisation angle

The pipeline offers potential integration into bioinformatic diagnostic software used by clinical laboratories and oncology research teams to interpret microarray gene expression data. The work represents an algorithmic framework tested on eight historical benchmark datasets, placing it at an early computational validation stage. Further clinical validation on prospective patient cohorts and compatibility testing with standard diagnostic workflows would be required before real-world adoption.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

Cancer prediction in the early stage is a topic of major interest in medicine since it allows accurate and efficient actions for successful medical treatments of cancer. Mostly cancer datasets contain various gene expression levels as features with less samples, so firstly there is a need to eliminate similar features to permit faster convergence rate of classification algorithms. These features (genes) enable us to identify cancer disease, choose the best prescription to prevent cancer and discover deviations amid different techniques. To resolve this problem, we proposed a hybrid novel technique CSSMO-based gene selection for cancer classification. First, we made alteration of the fitness of spider monkey optimization (SMO) with cuckoo search algorithm (CSA) algorithm viz., CSSMO for feature selection, which helps to combine the benefit of both metaheuristic algorithms to discover a subset of genes which helps to predict a cancer disease in early stage. Further, to enhance the accuracy of the CSSMO algorithm, we choose a cleaning process, minimum redundancy maximum relevance (mRMR) to lessen the gene expression of cancer datasets. Next, these subsets of genes are classified using deep learning (DL) to identify different groups or classes related to a particular cancer disease. Eight different benchmark microarray gene expression datasets of cancer have been utilized to analyze the performance of the proposed approach with different evaluation matrix such as recall, precision, F1-score, and confusion matrix. The proposed gene selection method with DL achieves much better classification accuracy than other existing DL and machine learning classification models with all large gene expression dataset of cancer.

Research topics

  • Gene expression and cancer classification
  • Machine Learning in Bioinformatics
  • Genetics, Bioinformatics, and Biomedical Research

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1186/s12859-023-05605-5

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.