Document Type


Date of Award

Fall 1-31-2002

Degree Name

Master of Science in Computer Science - (M.S.)


Computer and Information Science

First Advisor

Jason T. L. Wang

Second Advisor

Chengjun Liu

Third Advisor

Qicheng Ma


Data collection and cleaning is a very important part of an elaborate Data Mining System. 'TreeBASE' is a relational database of phylogenetic information at the Harvard University with a keyword based searching interface. 'TreeSearch' is a Structure based search engine implemented at NJIT that can be used for searching phylogenetic data. Phylogenetic trees are extracted from the flat-file database at Harvard University, available at {}. There is huge amount of information present in the files about the trees and the data matrices from which the trees are generated. The search tool implemented at NJIT is interested in using the string representation of the trees for query and retrieval of information.

The purpose of this thesis and the work related to it, is to make an automated tool to clean the files present in the Harvard University's Database using pattern-matching techniques and gather all the phylogenetic trees' strinng, representation to build a local database of clean phylogenetic data that can be readily used in a versatile fashion.



To view the content in your browser, please download Adobe Reader or, alternately,
you may Download the file to your hard drive.

NOTE: The latest versions of Adobe Reader do not support viewing PDF files within Firefox on Mac OS and if you are using a modern (Intel) Mac, there is no official plugin for viewing PDF files within the browser window.