Lossless Text Compression using Dictionaries
نویسنده
چکیده
Compression is used just about everywhere. Reduction of both compression ratio and retrieval of data from large collection is important in today‟s era. We propose a pre-compression technique that can be applied to text files. The output of our technique can be further applied to standard compression techniques available, such as arithmetic coding and BZIP2, which yields in better compression ratio. The algorithm suggested here uses the dynamic dictionary created at run-time and is also suitable for searching the phrases from the compressed file.
منابع مشابه
Efficient compression method for pronunciation dictionaries
Pronunciation dictionaries are often used with other datadriven methods to model the pronunciations in phonemebased automatic speech recognition (ASR) and text-to-speech (TTS) systems. The dictionaries usually take a great amount of memory, which is a limiting factor in portable handheld devices. Compressing the pronunciation dictionaries results in minimal transmission bandwidth and less stora...
متن کاملMorphological Analysis and Diacritical Arabic Text Compression
Morphological analysis of Arabic words allows decreasing the storage requirements of the Arabic dictionaries, more efficient encoding of diacritical Arabic text, faster spelling and efficient Optical character recognition. All these factors allow efficient storage and archival of multilingual digital libraries that include Arabic texts. This paper presents a lossless compression algorithm based...
متن کاملText Compression Algorithms - a Comparative Study
Data Compression may be defined as the science and art of the representation of information in a crisply condensed form. For decades, Data compression has been one of the critical enabling technologies for the ongoing digital multimedia revolution. There are a lot of data compression algorithms which are available to compress files of different formats. This paper provides a survey of different...
متن کاملDictionary-Based Fast Transform for Text Compression with High Compression Ratio
In this paper we introduce a dictionary-based fast lossless text transform algorithm. This algorithm utilizes ternary search tree to expedite transform encoding operation. Based on an efficient dictionary mapping model, this algorithm use a fast hash function to achieve a lightening speed in the transform decoding phrase. Results shows that the average compression time using the transform algor...
متن کاملDictionary-Based Fast Transform for Text Compression
In this paper we present StarNT, a dictionary-based fast lossless text transform algorithm. With a static generic dictionary, StarNT achieves a superior compression ratio than almost all the other recent efforts based on BWT and PPM. This algorithm utilizes ternary search tree to expedite transform encoding. Experimental results show that the average compression time has improved by orders of m...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2011