Your browser doesn't support javascript.
loading
Efficient sequencing data compression and FPGA acceleration based on a two-step framework.
Chen, Shifu; Chen, Yaru; Wang, Zhouyang; Qin, Wenjian; Zhang, Jing; Nand, Heera; Zhang, Jishuai; Li, Jun; Zhang, Xiaoni; Liang, Xiaoming; Xu, Mingyan.
Afiliação
  • Chen S; HaploX Biotechnology, Shenzhen, Guangdong, China.
  • Chen Y; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, Guangdong, China.
  • Wang Z; HaploX Biotechnology, Shenzhen, Guangdong, China.
  • Qin W; HaploX Biotechnology, Shenzhen, Guangdong, China.
  • Zhang J; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, Guangdong, China.
  • Nand H; HaploX Biotechnology, Shenzhen, Guangdong, China.
  • Zhang J; Xilinx Inc., San Jose, CA, United States.
  • Li J; Xilinx Inc., San Jose, CA, United States.
  • Zhang X; HaploX Biotechnology, Shenzhen, Guangdong, China.
  • Liang X; HaploX Biotechnology, Shenzhen, Guangdong, China.
  • Xu M; Xilinx Inc., San Jose, CA, United States.
Front Genet ; 14: 1260531, 2023.
Article em En | MEDLINE | ID: mdl-37811144
ABSTRACT
With the increasing throughput of modern sequencing instruments, the cost of storing and transmitting sequencing data has also increased dramatically. Although many tools have been developed to compress sequencing data, there is still a need to develop a compressor with a higher compression ratio. We present a two-step framework for compressing sequencing data in this paper. The first step is to repack original data into a binary stream, while the second step is to compress the stream with a LZMA encoder. We develop a new strategy to encode the original file into a LZMA highly compressed stream. In addition an FPGA-accelerated of LZMA was implemented to speedup the second step. As a demonstration, we present repaq as a lossless non-reference compressor of FASTQ format files. We introduced a multifile redundancy elimination method, which is very useful for compressing paired-end sequencing data. According to our test results, the compression ratio of repaq is much higher than other FASTQ compressors. For some deep sequencing data, the compression ratio of repaq can be higher than 25, almost four times of Gzip. The framework presented in this paper can also be applied to develop new tools for compressing other sequencing data. The open-source code of repaq is available at https//github.com/OpenGene/repaq.
Palavras-chave

Texto completo: 1 Coleções: 01-internacional Base de dados: MEDLINE Idioma: En Revista: Front Genet Ano de publicação: 2023 Tipo de documento: Article País de afiliação: China

Texto completo: 1 Coleções: 01-internacional Base de dados: MEDLINE Idioma: En Revista: Front Genet Ano de publicação: 2023 Tipo de documento: Article País de afiliação: China