Comparative Analysis of Clustering Methodologies in DNA Storage

  • Subhasiny Sankar
  • , Yixin Wang
  • , Zhang Jiayu
  • , Nur Sabrina
  • , Erry Gunawan
  • , Yong Liang Guan
  • , Noor A.Rahim Md
  • , Chueh Loo Poh

Research output: Chapter in Book/Report/Conference proceedingsChapterpeer-review

Abstract

Owing to the significance of DNA storage technology in meeting exponential storage demands and longevity, the challenges caused by bio-molecular errors while reading/sequencing data from DNA molecules must be addressed. By reading redundant copies, data can be reconstructed but with associated cost of sequencing and decoding complexities. Hence, solutions for dealing with both errors and complexities are sought after. The main objective of this work is to study data reconstruction methods for processing sequence readouts at downstream stage of DNA data storage. We investigated applicability of three clustering tools -Starcode, Slidesort, MeShClust, and two algorithms - Majority Nucleotide Selection (MNS), Cooperative Sequence Clustering (CSC) by transforming them into suitable tools for storage application. We observed that for fixed redundancy of 6.3x to 8.6x based on the nature of the dataset, Starcode outperforms other tools with 1% to 40% higher recovery rate. However, it costs the highest decoding complexity whereas MNS and CSC provides the lowest decoding complexity. Moreover, the distribution of the cluster and clustering speed of each tool/method are compared. This is the first comparative analysis study of tools/methods for data reconstruction in DNA data storage.

Original languageEnglish
Title of host publicationICSEC 2022 - International Computer Science and Engineering Conference 2022
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages269-274
Number of pages6
ISBN (Electronic)9781665491983
DOIs
Publication statusPublished - 2022
Event26th International Computer Science and Engineering Conference, ICSEC 2022 - Sakon Nakhon, Thailand
Duration: 21 Dec 202223 Dec 2022

Publication series

NameICSEC 2022 - International Computer Science and Engineering Conference 2022

Conference

Conference26th International Computer Science and Engineering Conference, ICSEC 2022
Country/TerritoryThailand
CitySakon Nakhon
Period21/12/2223/12/22

Keywords

  • Clustering
  • Data reconstruction
  • DNA data storage
  • Illumina sequencing

Fingerprint

Dive into the research topics of 'Comparative Analysis of Clustering Methodologies in DNA Storage'. Together they form a unique fingerprint.

Cite this