Clustering of Documents using Particle Swarm Optimization and Semantics Information
نویسندگان
چکیده
With the ever increasing volume of information, document clustering is used for automatic document organization so as to yield relevant information in an expeditious manner. Document clustering is an automatic grouping of text documents into clusters so that documents within a cluster have similar concepts. Representation of document is a very important step in any Information Retrieval (IR) system. In traditional document representation methods, the feature vector representing the document is constructed from the frequency count of document terms. But traditional document representation methods can not identify semantically related terms. In this paper, we present a semantic document clustering method that uses Universal Networking Language(UNL) and Particle Swarm Optimization(PSO). We generate feature vectors using UNL. The hybrid PSO+K-means algorithm is used to cluster the documents. Some experiments are performed to compare efficiency of the UNL method with the traditional term frequency based method. The results obtained show that the PSO-based clustering method using the UNL performs better than the term frequency based Method. Keywords—Universal Networking Language; Document clustering; Particle Swarm Optimization; K-means.
منابع مشابه
A Comparative Analysis of Particle Swarm Optimization and K-means Algorithm For Text Clustering Using Nepali Wordnet
The volume of digitized text documents on the web have been increasing rapidly. As there is huge collection of data on the web there is a need for grouping(clustering) the documents into clusters for speedy information retrieval. Clustering of documents is collection of documents into groups such that the documents within each group are similar to each other and not to documents of other groups...
متن کاملFuzzy clustering of time series data: A particle swarm optimization approach
With rapid development in information gathering technologies and access to large amounts of data, we always require methods for data analyzing and extracting useful information from large raw dataset and data mining is an important method for solving this problem. Clustering analysis as the most commonly used function of data mining, has attracted many researchers in computer science. Because o...
متن کاملFuzzy Particle Swarm Optimization Algorithm for a Supplier Clustering Problem
This paper presents a fuzzy decision-making approach to deal with a clustering supplier problem in a supply chain system. During recent years, determining suitable suppliers in the supply chain has become a key strategic consideration. However, the nature of these decisions is usually complex and unstructured. In general, many quantitative and qualitative factors, such as quality, price, and fl...
متن کاملStock Price Prediction using Machine Learning and Swarm Intelligence
Background and Objectives: Stock price prediction has become one of the interesting and also challenging topics for researchers in the past few years. Due to the non-linear nature of the time-series data of the stock prices, mathematical modeling approaches usually fail to yield acceptable results. Therefore, machine learning methods can be a promising solution to this problem. Methods: In this...
متن کاملAn Ontology Based Model for Document Clustering
Clustering is an important topic to find relevant content from a document collection and it also reduces the search space. The current clustering research emphasizes the development of a more efficient clustering method without considering the domain knowledge and user’s need. In recent years the semantics of documents have been utilized in document clustering. The discussed work focuses on the...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2014