<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD 2.3 20070202//EN" "http://dtd.nlm.nih.gov/publishing/2.3/journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
	<front>
		<journal-meta>
			<journal-id journal-id-type="nlm-ta">J Comput Sci Syst Biol</journal-id>
			<journal-id journal-id-type="publisher-id">opg</journal-id>						
			<journal-title>Journal of Computer Science &amp; Systems Biology</journal-title>			 
			<issn pub-type="epub">0974-7230</issn>
			<publisher>
				<publisher-name>OMICS Publishing Group</publisher-name>
				<publisher-loc>India, USA</publisher-loc>
			</publisher>
		</journal-meta>
		<article-meta>	
			<article-id pub-id-type="doi">10.4172/jcsb.1000031</article-id>		
			<article-id pub-id-type="publisher-id">000063</article-id>
			<article-categories>
				<subj-group subj-group-type="heading">
					<subject>Research Article</subject>
				</subj-group>
				<subj-group subj-group-type="Discipline">
					<subject>Biochemistry</subject>
				</subj-group>
				<subj-group subj-group-type="System Taxonomy">
					<subject>Proteomics</subject>
					<subject>Bioinformatics</subject>
					<subject>Genomics</subject>
					<subject>Transcriptomics</subject>
					<subject>Biomarkers</subject>
				</subj-group>
			</article-categories>
			<title-group>
				<article-title>Combination of Ant Colony Optimization and Bayesian Classification for Feature Selection in a Bioinformatics Dataset</article-title>
			</title-group>
			<contrib-group>
				<contrib contrib-type="author">
					<name>
						<surname>Hosseinzadeh Aghdam</surname>
						<given-names>Mehdi</given-names>						
					</name>
					<xref ref-type="aff" rid="a1">1</xref>
						<xref ref-type="corresp" rid="cor1">&ast;</xref>											
				</contrib>
				<contrib contrib-type="author">
					<name>
						<surname>Tanha</surname>
						<given-names>Jafar</given-names>
					</name>		
						<xref ref-type="aff" rid="a2">2</xref>										
				</contrib>
				<contrib contrib-type="author">
					<name>
						<surname>Naghsh-Nilchi</surname>
						<given-names>Ahmad Reza</given-names>
					</name>		
						<xref ref-type="aff" rid="a3">3</xref>										
				</contrib>
				<contrib contrib-type="author">
					<name>
						<surname>Ehsan Basiri</surname>
						<given-names>Mohammad</given-names>
					</name>		
						<xref ref-type="aff" rid="a3">3</xref>										
				</contrib>									
			</contrib-group>
			<aff id="a1"><label>1</label>Computer Engineering Department, Technical &amp; Engineering Faculty of Bonab, University of Tabriz, Tabriz, Iran</aff>
			<aff id="a2"><label>2</label>Computer Engineering Department &amp; IT, Payame Noor University, Lashgarak, MiniCity, Tehran, Iran</aff>
			<aff id="a3"><label>3</label>Computer Engineering Department, Faculty of Engineering, University of Isfahan, Hezar Jerib Avenue, Isfahan, Iran</aff>
			<author-notes>
				<corresp id="cor1">&ast; To whom correspondence should be addressed: Mehdi Hosseinzadeh Aghdam, Computer Engineering Department,Technical &amp; Engineering Faculty of Bonab, University of Tabriz, Tabriz, Iran; Phone:+98 311 7932671; Fax: +98 311 7932670; E-mail: <email>hosseinzadeh@comp.ui.ac.ir, mhaghdam@yahoo.com</email></corresp>
			</author-notes>
			<pub-date pub-type="collection">
			     <month>06</month>
				 <year>2009</year>
			</pub-date>
			<pub-date pub-type="epub">
				<day>15</day>
				<month>06</month>
				<year>2009</year>
			</pub-date>			
			<volume>2</volume>
			<issue>3</issue>
			<fpage>186</fpage>
			<lpage>199</lpage>
			<history>
			<date date-type="received">
			     <day>31</day>
				 <month>03</month>
				 <year>2009</year>
			</date>
			<date date-type="accepted">
			      <day>14</day>
				  <month>06</month>
				  <year>2009</year>
			</date>
			</history>
			<permissions>			 
			<copyright-statement><bold>Copyright:</bold> &copy; 2009 Aghdam MH, et al.</copyright-statement>
			<copyright-year>2009</copyright-year>
			<license license-type="open access">
			 <p>This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.</p>
			 </license>
			 </permissions>			
			<abstract>
				<p>Feature selection is widely used as the first stage of classification task to reduce the dimension of problem, decrease noise, improve speed and relieve memory constraints by the elimination of irrelevant or redundant features. One approach in the feature selection area is employing population-based optimization algorithms such as particle swarm optimization (PSO)-based method and ant colony optimization (ACO)-based method. Ant colony optimization algorithm is inspired by observation on real ants in their search for the shortest paths to food sources. Protein function prediction is an important problem in functional genomics. Typically, protein sequences are represented by feature vectors. A major problem of protein datasets that increase the complexity of classification models is their large number of features. This paper empowers the ant colony optimization algorithm by enabling the ACO to select features for a Bayesian classification method. The naive Bayesian classifier is a straightforward and frequently used method for supervised learning. It provides a flexible way for dealing with any number of features or classes, and is based on probability theory. This paper then compares the performance of the proposed ACO algorithm against the performance of a standard binary particle swarm optimization algorithm on the task of selecting features on Postsynaptic dataset. The criteria used for this comparison are maximizing predictive accuracy and finding the smallest subset of features. Simulation results on Postsynaptic dataset show that proposed method simplifies features effectively and obtains a higher classification accuracy compared to other feature selection methods.</p>
			</abstract>
			 <kwd-group>
				<kwd>Feature Selection</kwd>
				<kwd>Ant Colony Optimization</kwd>
				<kwd>Particle Swarm Optimization</kwd>
				<kwd>Bayesian Classification</kwd>
				<kwd>Bioinformatics</kwd>										
			</kwd-group>
			<custom-meta-wrap>
				<custom-meta>
					<meta-name>citation</meta-name>
					<meta-value>Aghdam MH, Tanha J, Naghsh-Nilchi AR, Basiri ME (2009) Combination of Ant Colony Optimization and Bayesian Classification for Feature Selection in a Bioinformatics Dataset. J Comput Sci Syst Biol 2: 186-199. doi:<ext-link ext-link-type="doi" xlink:href="10.4172/jcsb.1000031">10.4172/jcsb.1000031</ext-link></meta-value>
				</custom-meta>
			</custom-meta-wrap>
		</article-meta>
	</front>	
	<body>
	     <sec>
		     <title>Introduction</title>
			 <p>The feature selection problem can be viewed as a particular case of a more general subset selection problem in which the goal is to find a subset maximizing some adopted criterion. Feature selection methods search through the subsets of features and try to find the best subset among the competing 2<sup><italic>N</italic></sup>-1 candidate subsets according to some evaluation measure, where N denotes the total number of features. Feature selection is used in many application areas as a tool to remove irrelevant and redundant features. The objective of feature selection is to simplify a dataset by reducing its dimensionality and identifying relevant underlying features without sacrificing predictive accuracy (<xref ref-type="bibr" rid="r17">Jensen, 2005</xref>). In many applications, the size of a dataset is so large that learning might not work as well before removing these unwanted features. Reducing the number of irrelevant/redundant features drastically reduces the running time of a learning algorithm and yields a more general concept. This helps in getting a better insight into the underlying concept of a real world classification problem (<xref ref-type="bibr" rid="r6">Dash &amp; Liu, 1997</xref>).</p>
			 <p>Protein function prediction is an important problem in functional genomics. Proteins are large molecules that perform nearly all of the functions of a cell in a living organism (<xref ref-type="bibr" rid="r2">Alberts et al., 2002</xref>). The primary sequence of a protein consists of a long sequence of amino acids. Proteins are the most essential and versatile macromolecules of life, and the knowledge of their functions is a crucial link in the development of new drugs, better crops, and even the development of synthetic biochemical such as biofuels.</p>
			 <p>Over the past few decades, major advances in the field of molecular biology, coupled with advances in genomic technologies, have led to an explosive growth in the biological information generated by the scientific community. Although the number of proteins with known sequence has grown exponentially in the last few years, due to rapid advances in genome sequencing technology, the number of proteins with known structure and function has grown at a substantially lower rate (<xref ref-type="bibr" rid="r14">Freitas &amp; de Carvalho, 2007</xref>).</p>
			 <p>Searching for similar sequences in protein databases is a common approach used in the prediction of a protein function. The objective of this search is to find a similar protein whose function is known and assigning its function to the new protein. Despite the simplicity and usefulness this method in a large number of situations, it has also some limitations (<xref ref-type="bibr" rid="r14">Freitas &amp; de Carvalho, 2007</xref>). For instance, two proteins might have very similar sequences and perform different functions, or have very different sequences and perform a similar function. Additionally, the proteins being compared may be similar in regions of the sequence that are not determinants of their function.</p>
			 <p>Another approach that may be used alternatively or in complement to the similarity-based approach is to build a model for predictive classification. The goal of such a model is to classify data instances into one of a predefined set of classes or categories. In this approach a feature vector represents each protein, a learning algorithm captures the most important relationships between the features, and the classes present in the dataset. A major problem in protein datasets is the high dimensionality of the feature space. Most of these dimensions are not relative to protein function; even some noise data hurt the performance of the classifier. Hence, we need to select some representative features from the original feature space to reduce the dimensionality of feature space and improve the efficiency and performance of classifier.</p>
			 <p>Among too many methods which are proposed for feature selection, population-based optimization algorithms such as particle swarm optimization (PSO)-based method (<xref ref-type="bibr" rid="r31">Wang, 2007</xref>) and ant colony optimization (ACO)-based method (Aghdam et al., 2008) have attracted a lot of attention. These methods attempt to achieve better solutions by application of knowledge from previous iterations.</p>
			 <p>Particle swarm optimization comprises a set of search techniques, inspired by the behavior of natural swarms, for solving optimization problems (<xref ref-type="bibr" rid="r19">Kennedy &amp; Eberhart, 2001</xref>). PSO is a global optimization algorithm for dealing with problems in which a point or surface in an n-dimensional space best represents a solution. Potential solutions are plotted in this space and seeded with an initial velocity. Particles move through the solution space and certain fitness criteria evaluate them. After a while particles accelerate toward those with better fitness values.</p>
			 <p>Meta-heuristic optimization algorithm based on ant's behavior was represented in the early 1990s by M. Dorigo and colleagues (<xref ref-type="bibr" rid="r9">Dorigo &amp; Caro, 1999</xref>). Ant colony optimization is a branch of newly developed form of artificial intelligence called swarm intelligence. Swarm intelligence is a field which studies "the emergent collective intelligence of groups of simple agents" (<xref ref-type="bibr" rid="r4">Bonabeau et al., 1999</xref>). In groups of insects which live in colonies, such as ants and bees, an individual can only do simple task on its own, while the colony's cooperative work is the main reason determining the intelligent behavior it shows (<xref ref-type="bibr" rid="r22">Liu et al., 2004</xref>).</p>
			 <p>ACO algorithm is inspired by ant's social behavior. Ants have no sight and are capable of finding the shortest route between a food source and their nest by chemical materials called pheromone that they leave when moving (<xref ref-type="bibr" rid="r4">Bonabeau et al., 1999</xref>). ACO algorithm was firstly used for solving traveling salesman problem (TSP) (<xref ref-type="bibr" rid="r10">Dorigo et al., 1996</xref>) and then has been successfully applied to a large number of difficult problems like the quadratic assignment problem (QAP) (<xref ref-type="bibr" rid="r24">Maniezzo &amp; Colorni, 1999</xref>), routing in telecommunication networks, graph coloring problems, scheduling, etc. This method is particularly attractive for feature selection as there seems to be no heuristic that can guide search to the optimal minimal subset every time (<xref ref-type="bibr" rid="r17">Jensen, 2005</xref>). On the other hand, if features are represented as a graph, ants will discover best feature combinations as they traverse the graph.</p>
			 <p>This paper proposed an ACO-based algorithm for the feature selection task in bioinformatics datasets. This paper extended by taking advantage of naive Bayes classifier and enable ACO to select features for a Bayesian classification method, which is more sophisticated than the nearest neighbor classifier. Bayesian classifiers are statistical classifiers that can predict class membership probabilities, such as the probability that a given sample belongs to a particular class. Bayesian classification is based on the Bayes theorem (<xref ref-type="bibr" rid="r12">Feller, 1971</xref>).</p>
			 <p>The naive Bayes (NB) classifier uses a probabilistic approach to assign each record of the dataset to a possible class. A naive Bayes classifier makes significant use of the assumption that all features are conditionally independent of one another given the class. This assumption is called class conditional independence (<xref ref-type="bibr" rid="r25">Mitchell, 1996</xref>). It is made to simplify the computations involved and, in this sense, is considered naive. In practice, dependencies can exist between variables; however, when the assumption holds true, then the Naive Bayes classifier is the most accurate in comparison with other classifiers (<xref ref-type="bibr" rid="r16">Han &amp; Kamber, 2001</xref>).</p>
			 <p>For testing the proposed ACO algorithm, it is applied to the problem of predicting whether or not a protein has a post-synaptic activity, based on features of protein's primary sequence and finally, the classifier performance and the length of selected feature subset are considered for performance evaluation.</p>
			 <p>The rest of this paper is organized as follows. Section 2 presents a brief overview of feature selection methods. Sections 3 and 4 briefly address K-nearest neighbor and naive Bayes classifiers. Ant colony optimization is described in sections 5. Section 6 explains the particle swarm optimization. Section 7 reports computational experiments. It also includes a brief discussion of the results obtained and finally the conclusion is offered in the last section.</p>
			 </sec>
			 <sec>
			     <title>Feature Selection Approaches</title>
				 <p>Feature selection is a process that selects a subset of original features. The optimality of a feature subset is measured by an evaluation criterion. As the dimensionality of a domain expands, the number of features increases. Finding an optimal feature subset is usually intractable (<xref ref-type="bibr" rid="r20">Kohavi &amp; John, 1997</xref>) and many problems related to feature selection have been shown to be NP-hard. A typical feature selection process consists of four basic steps, namely, subset generation, subset evaluation, stopping criterion, and result validation (<xref ref-type="bibr" rid="r6">Dash &amp; Liu, 1997</xref>). Subset generation is a search procedure that produces candidate feature subsets for evaluation based on a certain <italic>search strategy</italic>. Each candidate subset is evaluated and compared with the previous best one according to a certain <italic>evaluation criterion</italic>. If the new subset turns out to be better, it replaces the previous best subset. The process of subset generation and evaluation is repeated until a given stopping criterion is satisfied. Then the selected best subset usually needs to be validated by prior knowledge or different tests via synthetic and/or real world datasets.</p>
				 <p>The generation procedure implements a search method (<xref ref-type="bibr" rid="r30">Siedlecki &amp; Sklansky, 1988</xref>) that generates subsets of features for evaluation. It may start with no features, all features, a selected feature set or some random feature subset. Those methods that start with an initial subset usually select these features heuristically beforehand. Features are added (<italic>forward selection</italic>) or removed (<italic>backward elimination</italic>) iteratively in the first two cases (<xref ref-type="bibr" rid="r6">Dash &amp; Liu, 1997</xref>). In the last case, features are either iteratively added or removed or produced randomly thereafter (<xref ref-type="bibr" rid="r17">Jensen, 2005</xref>). The disadvantage of forward selection and backward elimination methods is that the features that were once selected/eliminated cannot be later discarded/re-selected. To overcome this problem, Pudil et al. proposed a method to flexibly add and remove features (<xref ref-type="bibr" rid="r28">Pudil et al., 1994</xref>). This method has been called floating search method.</p>
				 <p>According to the literature, the approaches to feature subset selection can be divided into filters and wrappers approaches (<xref ref-type="bibr" rid="r15">Guyon &amp; Elisseeff, 2003</xref>). The filter model separates feature selection from classifier learning and selects feature subsets that are independent of any learning algorithm. It relies on various measures of the general characteristics of the training data such as distance, information, dependency, and consistency (<xref ref-type="bibr" rid="r23">Liu &amp; Motoda, 1998</xref>). In the wrapper approach feature subset is selected using the evaluation function based on the same learning algorithm that will be used later for learning. In this approach the evaluation function calculates the suitability of a feature subset produced by the generation procedure and it also compares that with the previous best candidate, replacing it if found to be better. A stopping criterion is tested in each of iterations to determine whether or not the feature selection process should continue. Although, wrappers may produce better results, they are expensive to run and can break down with very large numbers of features. This is due to the use of learning algorithms in the evaluation of subsets, some of which can encounter problems while dealing with large datasets (<xref ref-type="bibr" rid="r13">Forman, 2003</xref>; <xref ref-type="bibr" rid="r17">Jensen, 2005</xref>).</p>				 
			 <sec>
			     <title>Literature Review</title>
				 <p>John, Kohavi and Pfleger addressed the problem of irrelevant features and the subset selection problem. They presented definitions for irrelevance and for two degrees of relevance (weak and strong). They also state that features selected should depend not only on the features and the target concept, but also on the induction algorithm (<xref ref-type="bibr" rid="r18">John et al., 1994</xref>).</p>
				 <p>Pudil, Novovicova and Kittler presented "floating" search methods in feature selection. These are sequential search methods characterized by a dynamically changing number of features included or eliminated at each step. They were shown to give very good results and to be computationally more effective than the branch and bound method (<xref ref-type="bibr" rid="r28">Pudil et al., 1994</xref>).</p>
				 <p>Dash and Liu gave a survey of feature selection methods for classification (<xref ref-type="bibr" rid="r6">Dash &amp; Liu, 1997</xref>).</p>
				 <p>Kohavi and John introduced wrappers for feature subset selection. Their approach searches for an optimal feature subset tailored to a particular learning algorithm and a particular training set (<xref ref-type="bibr" rid="r20">Kohavi &amp; John, 1997</xref>).</p>
				 <p>Yang and Honavar used a genetic algorithm for feature subset selection (<xref ref-type="bibr" rid="r33">Yang &amp; Honavar, 1998</xref>).</p>
				 <p>Liu and Motoda wrote their book on feature selection which offers an overview of the methods developed since the 1970s and provides a general framework in order to examine these methods and categorize them (<xref ref-type="bibr" rid="r23">Liu &amp; Motoda, 1998</xref>).</p>
				 <p>Forman presented an empirical comparison of twelve feature selection methods. Results revealed the surprising performance of a new feature selection metric, 'Bi-Normal Separation' (BNS) (<xref ref-type="bibr" rid="r13">Forman, 2003</xref>).</p>
				 <p>Guyon and Elisseeff gave an introduction to variable and feature selection. They recommend using a linear predictor of your choice (e.g. a linear SVM) and select variables in two alternate ways: (1) with a variable ranking method using correlation coefficient or mutual information; (2) with a nested subset selection method performing forward or backward selection or with multiplicative updates (<xref ref-type="bibr" rid="r15">Guyon &amp; Elisseeff, 2003</xref>).</p>
			 </sec>
			</sec>
			<sec>
			     <title><italic>K</italic>-Nearest Neighbor Classifier</title>
				 <p>The <italic>K</italic>-nearest neighbor (KNN) algorithm is amongst the simplest of all machine learning algorithms. An object is classified by a majority vote of its neighbors, with the object being assigned to the class most common amongst its <italic>K</italic> nearest neighbors. <italic>K</italic> is a positive integer, typically small. If <italic>K</italic> = 1, then the object is simply assigned to the class of its nearest neighbor. In binary (two class) classification problems, it is helpful to choose <italic>K</italic> to be an odd number as this avoids tied votes.</p>
				 <p>The same method can be used for regression, by simply assigning the property value for the object to be the average of the values of its <italic>K</italic> nearest neighbors. It can be useful to weight the contributions of the neighbors, so that the nearer neighbors contribute more to the average than the more distant ones.</p>
				 <p>The neighbors are taken from a set of objects for which the correct classification (or, in the case of regression, the value of the property) is known. This can be thought of as the training set for the algorithm, though no explicit training step is required. In order to identify neighbors, the objects are represented by position vectors in a multidimensional feature space. It is usual to use the Euclidean distance, though other distance measures, such as the Manhattan distance could in principle be used instead. The <italic>K</italic>-nearest neighbor algorithm is sensitive to the local structure of the data.</p>
				 <p>The performance of a KNN classifier is primarily determined by the choice of <italic>K</italic> as well as the distance metric applied (<xref ref-type="bibr" rid="r21">Latourrette, 2000</xref>). However, it has been shown in (<xref ref-type="bibr" rid="r7">Domeniconi et al., 2002</xref>) that when the points are not uniformly distributed, predetermining the value of <italic>K</italic> becomes difficult. Generally, larger values of <italic>K</italic> are more immune to the noise presented, and make boundaries smoother between classes. As a result, choosing the same (optimal) <italic>K</italic> becomes almost impossible for different applications.</p>
			 </sec>
			 <sec>
			     <title>Naive Bayes Classifier</title>
				 <p>One highly practical Bayesian learning method is the naive Bayes leaner, often called the naive Bayes (NB) classifier. In some domains its performance has been shown to be comparable to that of neural network and decision tree learning. This section introduces the naive Bayes classifier.</p>
				 <p>The naive Bayes classifier applies to classification tasks where each instance x is described by a conjunction of feature values and where the target function f(x) can take on any value from some finite set <italic>V</italic>. A set of training examples of the target function is provided, and a new instance is presented, described by the tuple of feature values &lt;<italic>a<sub>1</sub>,a<sub>2</sub>...a<sub>n</sub></italic>&gt;. The classifier is asked to predict the target value, or classification, for this new instance (<xref ref-type="bibr" rid="r25">Mitchell, 1996</xref>).</p>
				 <p>The Bayesian approach to classifying the new instance is to assign the most probable target value, <italic>v<sub>MAP</sub></italic>, given the feature values &lt;a<sub>1</sub>,a<sub>2</sub>...a<sub>n</sub>&gt; that describe in instance.</p>
				 <p><disp-formula id="E1"><label>[1]</label><mml:math id="M1" display='block'> <mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munder><mml:mrow><mml:mi>arg</mml:mi><mml:mi>max</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x007C;</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mn>2</mml:mn></mml:msub> <mml:mn>...</mml:mn><mml:msub><mml:mi>a</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
				 <p>We can use Bayes theorem to rewrite this expression as</p>
				 <p><disp-formula id="E2"><label>[2]</label><mml:math id="M2" display='block'><mml:mtable columnalign='left'>
  <mml:mtr>
   <mml:mtd>
    <mml:msub>
     <mml:mi>v</mml:mi>
     <mml:mrow>
      <mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>P</mml:mi>
     </mml:mrow>
    </mml:msub>
    <mml:mo>=</mml:mo><mml:munder>
     <mml:mrow>
      <mml:mi>arg</mml:mi><mml:mi>max</mml:mi>
     </mml:mrow>
     <mml:mrow>
      <mml:msub>
       <mml:mrow></mml:mrow>
       <mml:mrow>
        <mml:msub>
         <mml:mi>v</mml:mi>
         <mml:mi>j</mml:mi>
        </mml:msub>
        <mml:mo>&#x2208;</mml:mo><mml:mi>V</mml:mi>
       </mml:mrow>
      </mml:msub>     
     </mml:mrow>
    </mml:munder>
    <mml:mfrac>
     <mml:mrow>
      <mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub>
       <mml:mi>a</mml:mi>
       <mml:mn>1</mml:mn>
      </mml:msub>
      <mml:mo>,</mml:mo><mml:msub>
       <mml:mi>a</mml:mi>
       <mml:mn>2</mml:mn>
      </mml:msub>
      <mml:mn>...</mml:mn><mml:msub>
       <mml:mi>a</mml:mi>
       <mml:mi>n</mml:mi>
      </mml:msub>
      <mml:mo>&#x007C;</mml:mo><mml:msub>
       <mml:mi>v</mml:mi>
       <mml:mi>j</mml:mi>
      </mml:msub>
      <mml:mo stretchy='false'>)</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub>
       <mml:mi>v</mml:mi>
       <mml:mi>j</mml:mi>
      </mml:msub>
      <mml:mo stretchy='false'>)</mml:mo>
     </mml:mrow>
     <mml:mrow>
      <mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub>
       <mml:mi>a</mml:mi>
       <mml:mn>1</mml:mn>
      </mml:msub>
      <mml:mo>,</mml:mo><mml:msub>
       <mml:mi>a</mml:mi>
       <mml:mn>2</mml:mn>
      </mml:msub>
      <mml:mn>...</mml:mn><mml:msub>
       <mml:mi>a</mml:mi>
       <mml:mi>n</mml:mi>
      </mml:msub>
      <mml:mo stretchy='false'>)</mml:mo>
     </mml:mrow>
    </mml:mfrac>   
   </mml:mtd>
  </mml:mtr>
  <mml:mtr>
   <mml:mtd>
    <mml:mrow></mml:mrow>
   </mml:mtd>
  </mml:mtr>
  <mml:mtr>
   <mml:mtd>
    <mml:mtext>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</mml:mtext><mml:mo>=</mml:mo><mml:munder>
     <mml:mrow>
      <mml:mi>arg</mml:mi><mml:mi>max</mml:mi>
     </mml:mrow>
     <mml:mrow>
      <mml:msub>
       <mml:mrow></mml:mrow>
       <mml:mrow>
        <mml:msub>
         <mml:mi>v</mml:mi>
         <mml:mi>j</mml:mi>
        </mml:msub>
        <mml:mo>&#x2208;</mml:mo><mml:mi>V</mml:mi>
       </mml:mrow>
      </mml:msub>      
     </mml:mrow>
    </mml:munder>
    <mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub>
     <mml:mi>a</mml:mi>
     <mml:mn>1</mml:mn>
    </mml:msub>
    <mml:mo>,</mml:mo><mml:msub>
     <mml:mi>a</mml:mi>
     <mml:mn>2</mml:mn>
    </mml:msub>
    <mml:mn>...</mml:mn><mml:msub>
     <mml:mi>a</mml:mi>
     <mml:mi>n</mml:mi>
    </mml:msub>
    <mml:mo>&#x007C;</mml:mo><mml:msub>
     <mml:mi>v</mml:mi>
     <mml:mi>j</mml:mi>
    </mml:msub>
    <mml:mo stretchy='false'>)</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub>
     <mml:mi>v</mml:mi>
     <mml:mi>j</mml:mi>
    </mml:msub>
    <mml:mo stretchy='false'>)</mml:mo>
   </mml:mtd>
  </mml:mtr>
 </mml:mtable> 
</mml:math>
</disp-formula>
</p>
<p>We could attempt to estimate the two terms in equation (2) based on the training data. It is easy to estimate each of the <italic>P</italic>(<italic>v<sub>i</sub></italic>) simply by counting the frequency with which each target value <italic>v<sub>i</sub></italic> occurs in the training data. However, estimating the different <italic>P</italic>(<italic>a</italic><sub>1</sub>,<italic>a</italic><sub>2</sub>...<italic>a</italic><sub>n</sub> | <italic>v<sub>j</sub></italic>) terms in this fashion is not feasible unless we have a very, very large set of training data. The problem is that the number of these terms is equal to the number of possible instances times the number of possible target values. Therefore, we need to see every instance in the instance space many times in order to obtain reliable estimates.</p>
<p>The naive Bayes classifier is based on the simplifying assumption that the feature values are conditionally independent given the target value. In other words, the assumption is that given the target value of the instance, the probability of observing the conjunction <italic>a<sub>1</sub>,a<sub>2</sub>...a<sub>n</sub></italic> is just the product of the probabilities for the individual features: <italic>P</italic>(<italic>a<sub>1</sub>,a<sub>2</sub>...a<sub>1</sub> | v<sub>j</sub></italic>) = <italic>P</italic>(<italic>v<sub>j</sub></italic>)&Pi;<sub>i</sub><italic>P</italic>(<italic>a<sub>i</sub> | v<sub>j</sub></italic>). Substituting this into equation (2), we have the approach used by the naive Bayes classifier.</p>
<p><disp-formula id="E3"><label>[3]</label><mml:math id="M3" display='block'>
 <mml:mrow>
  <mml:msub>
   <mml:mi>v</mml:mi>
   <mml:mrow>
    <mml:mi>N</mml:mi><mml:mi>B</mml:mi>
   </mml:mrow>
  </mml:msub>
  <mml:mo>=</mml:mo><mml:munder>
   <mml:mrow>
    <mml:mi>arg</mml:mi><mml:mi>max</mml:mi>
   </mml:mrow>
   <mml:mrow>
    <mml:msub>
     <mml:mrow></mml:mrow>
     <mml:mrow>
      <mml:msub>
       <mml:mi>v</mml:mi>
       <mml:mi>j</mml:mi>
      </mml:msub>
      <mml:mo>&#x2208;</mml:mo><mml:mi>V</mml:mi>
     </mml:mrow>
    </mml:msub>
    
   </mml:mrow>
  </mml:munder>
  <mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub>
   <mml:mi>v</mml:mi>
   <mml:mi>j</mml:mi>
  </mml:msub>
  <mml:mo stretchy='false'>)</mml:mo><mml:munder>
   <mml:mstyle mathsize='140%' displaystyle='true'><mml:mi>&#x03A0;</mml:mi></mml:mstyle>
   <mml:mi>i</mml:mi>
  </mml:munder>
  <mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub>
   <mml:mi>a</mml:mi>
   <mml:mi>i</mml:mi>
  </mml:msub>
  <mml:mo>&#x007C;</mml:mo><mml:msub>
   <mml:mi>v</mml:mi>
   <mml:mi>j</mml:mi>
  </mml:msub>
  <mml:mo stretchy='false'>)</mml:mo>
 </mml:mrow>
</mml:math>
</disp-formula>
</p>
<p>where <italic>v<sub>NB</sub></italic> denotes the target value output by the naive Bayes classifier (<xref ref-type="bibr" rid="r25">Mitchell, 1996</xref>).</p>
<p>One interesting difference between the naive Bayes classification method and other classification methods we have considered is that there is no explicit search through the space of possible hypotheses (in this case, the space of possible hypotheses is the space of possible values that can be assigned to the various <italic>P</italic>(<italic>v<sub>i</sub></italic>) and <italic>P</italic>(<italic>a<sub>i</sub></italic> | <italic>v<sub>i</sub></italic>) terms). Instead, the hypothesis is formed without searching, simply by counting the frequency of various data combinations within the training examples.</p>
<p>To evaluate the performance of the Bayesan classification, we use a 10-fold cross-validation. We divide the data set into 10 equally sized folds. For all class levels each fold maintains roughly the same proportion of classes present in the whole data set before division (called stratified crossvalidation). Eight of the ten folds are used to compute the probabilities for the Bayesian classification. The ninth fold is used as validation set and the tenth fold as test set. During the search for the solution only the validation set is used to compute predictive accuracy. The performance of the candidate solutions is given by the predictive accuracy of the classification in the validation set. The solution that shows the highest predictive accuracy on the validation set is then used to compute the predictive accuracy on the test set. Once the solution is selected, the nine folds are merged and this merged dataset is used to compute the probabilities for the Bayesian classification. The predictive accuracy (reported as the final result) is then computed on the previously untouched test set fold. Every fold will be once used as validation set and once used as test set. A similar process is adopted for the computation of the predictive accuracy using the nearest neighbor classifier, which will be described in more details in subsection 7.2.</p>
			 </sec>
		 <sec>
			     <title>Ant Colony Optimization</title>
				 <p>Ant colony optimization was introduced by Marco Dorigo (<xref ref-type="bibr" rid="r8">Dorigo, 1992</xref>) and his colleagues in the early 1990s. The first computational paradigm appeared under the name ant system (AS). It is another approach to stochastic combinatorial optimization. The search activities are distributed over: "ants" - agents with very simple basic capabilities that mimic the behavior of real ants. The main aim was not to simulate ant colonies, but to use artificial ant colonies as an optimization tool. Therefore the system exhibits several differences in comparison to the real (natural) ant colony: artificial ants have some memory; they are not completely blind; they live in an environment with discrete time. In ACO algorithms, artificial ants construct solutions from scratch by probabilistically making a sequence of local decisions. At each construction step an ant chooses exactly one of possibly several ways of extending the current partial solution. The rules that define the solution construction mechanism in ACO implicitly map the search space of the considered problem (including the partial solutions) onto a search tree.</p>
				 <p>The paradigm is based on the observation made by ethologists about the medium used by ants to communicate information regarding shortest paths to food by means of pheromone trails. A moving ant lays some pheromone on the ground, thus making a path by a trail of this substance. While an isolated ant moves practically at random (exploration), an ant encountering a previously laid trail can detect it and decide with high probability to follow it and consequently reinforce the trail with its own pheromone (exploitation). What emerges is a form of autocatalytic process through which the more the ants follow a trail, the more attractive that trail becomes to be followed. The process is thus characterized by a positive feedback loop, during which the probability of choosing a path increases with the number of ants that previously chose the same path. The mechanism above is the inspiration for the algorithms of the ACO family (<xref ref-type="bibr" rid="r11">Engelbrecht, 2005</xref>).</p>
			 </sec>
			 <sec>
			     <title>Proposed ACO Algorithm for Feature Selection</title>
				 <p>Feature selection is one of the applications of subset problems. Given an feature set of size n, the feature selection problem is to find a minimal feature subset of size s (s &lt; n) while retaining a suitably high accuracy in representing the original features. Therefore, there is no concept of path. A partial solution does not define any ordering among the components of the solution, and the next component to be selected is not necessarily influenced by the last component added to the partial solution (<xref ref-type="bibr" rid="r1">Aghdam, 2008</xref>). Furthermore, solutions to a feature selection problem are not necessarily of the same size. To apply an ACO algorithm to solve a feature selection problem, these aspects need to be addressed. The first problem is addressed by redefining the way that the representation graph is used.</p>
				 <p>The feature selection task may be reformulated into an ACO-suitable problem. Ant colony optimization requires a problem to be represented as a graph. Here nodes represent features, with the edges between them denoting the choice of the next feature. The search for the optimal feature subset is then an ant traversal through the graph where a minimum number of nodes are visited that satisfies the traversal stopping criterion. <xref ref-type="fig" rid="g1">Figure 1.</xref> illustrates this setup. The ant is currently at node <italic>a</italic> and has a choice of which feature to add next to its path (dotted lines). It chooses feature <italic>b</italic> next based on the transition rule, then <italic>c</italic> and then <italic>d</italic>. Upon arrival at <italic>d</italic>, the current subset {<italic>a, b, c, d</italic>} is determined to satisfy the traversal stopping criterion (e.g. suitably high classification accuracy has been achieved with this subset). The ant terminates its traversal and outputs this feature subset as a candidate for data reduction (<xref ref-type="bibr" rid="r1">Aghdam, 2008</xref>).</p>
				 <p>A suitable heuristic desirability of traversing between features could be any subset evaluation function for example, an entropy-based measure (<xref ref-type="bibr" rid="r17">Jensen, 2005</xref>) or rough set dependency measure (<xref ref-type="bibr" rid="r27">Pawlak, 1991</xref>). The heuristic desirability of traversal and edge pheromone levels are combined to form the so-called probabilistic transition rule, denoting the probability that ant k will include feature i in its solution at time step t:</p>
				 <p><disp-formula id="E4"><label>[4]</label><mml:math id="M4" display='block'><mml:mrow>
  <mml:msubsup>
   <mml:mstyle mathsize='140%' displaystyle='true'><mml:mi>P</mml:mi></mml:mstyle>
   <mml:mi>i</mml:mi>
   <mml:mi>k</mml:mi>
  </mml:msubsup>
  <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo> <mml:mtable columnalign='left'>
   <mml:mtr>
    <mml:mtd>
     <mml:mfrac>
      <mml:mrow>
       <mml:msup>
        <mml:mrow>
         <mml:mo stretchy='false'>[</mml:mo><mml:msub>
          <mml:mi>&#x03C4;</mml:mi>
          <mml:mi>i</mml:mi>
         </mml:msub>
         <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>]</mml:mo>
        </mml:mrow>
        <mml:mi>&#x03B1;</mml:mi>
       </mml:msup>
       <mml:mo>.</mml:mo><mml:msup>
        <mml:mrow>
         <mml:mo stretchy='false'>[</mml:mo><mml:msub>
          <mml:mi>&#x03B7;</mml:mi>
          <mml:mi>i</mml:mi>
         </mml:msub>
         <mml:mo stretchy='false'>]</mml:mo>
        </mml:mrow>
        <mml:mi>&#x03B2;</mml:mi>
       </mml:msup>       
      </mml:mrow>
      <mml:mrow>
       <mml:mstyle displaystyle='true'>
        <mml:munder>
         <mml:mo>&#x2211;</mml:mo>
         <mml:mrow>
          <mml:mi>l</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup>
           <mml:mi>J</mml:mi>
           <mml:mi>k</mml:mi>
          </mml:msup>          
         </mml:mrow>
        </mml:munder>
        <mml:mrow>
         <mml:msup>
          <mml:mrow>
           <mml:mo stretchy='false'>[</mml:mo><mml:msub>
            <mml:mi>&#x03C4;</mml:mi>
            <mml:mi>l</mml:mi>
           </mml:msub>
           <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>]</mml:mo>
          </mml:mrow>
          <mml:mi>&#x03B1;</mml:mi>
         </mml:msup>
         <mml:mo>.</mml:mo><mml:msup>
          <mml:mrow>
           <mml:mo stretchy='false'>[</mml:mo><mml:msub>
            <mml:mi>&#x03B7;</mml:mi>
            <mml:mi>l</mml:mi>
           </mml:msub>
           <mml:mo stretchy='false'>]</mml:mo>
          </mml:mrow>
          <mml:mi>&#x03B2;</mml:mi>
         </mml:msup>         
        </mml:mrow>
       </mml:mstyle>
      </mml:mrow>
     </mml:mfrac>
     <mml:mtext>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</mml:mtext><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mtext>&nbsp;</mml:mtext><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup>
      <mml:mi>J</mml:mi>
      <mml:mi>k</mml:mi>
     </mml:msup>     
    </mml:mtd>
   </mml:mtr>
   <mml:mtr>
    <mml:mtd>
     <mml:mn>0</mml:mn><mml:mtext>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</mml:mtext><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi>
    </mml:mtd>
   </mml:mtr>
  </mml:mtable>
   </mml:mrow>
 </mml:mrow>
</mml:math>
</disp-formula>
</p>
<p>where <italic>&eta;<sub>j</sub></italic> is the heuristic desirability of choosing feature i (<italic>&eta;<sub>j</sub></italic>is optional but often needed for achieving a high algorithm performance), <italic>J<sup>k</sup></italic> is the set of feasible features that can be added to the partial solution. <italic>&alpha;</italic> &gt; 0, <italic>&beta;</italic> &gt; 0 are two parameters that determine the relative importance of the pheromone value and heuristic information (the choice of <italic>&alpha;</italic>, <italic>&beta;</italic> is determined experimentally) and <italic>&tau;<sub>j</sub></italic>(<italic>t</italic>) is the amount of virtual pheromone on feature <italic>i</italic>.</p>
<p>                 
                    <fig id="g1">
					<label>Figure 1</label>
					<caption>
						<title>ACO problem representation for feature selection.</title>						
					</caption>
					<graphic xlink:href="JCSB-02-186-g001.tif"/>
				</fig>
				</p>
				<p>The pheromone on each feature is updated according to the following formula:</p>
				<p><disp-formula id="E5"><label>[5]</label><mml:math id="M5" display='block'>
 <mml:mrow>
  <mml:msub>
   <mml:mi>&#x03C4;</mml:mi>
   <mml:mi>i</mml:mi>
  </mml:msub>
  <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>&#x03C1;</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>.</mml:mo><mml:msub>
   <mml:mi>&#x03C4;</mml:mi>
   <mml:mi>i</mml:mi>
  </mml:msub>
  <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mstyle displaystyle='true'>
   <mml:munderover>
    <mml:mo>&#x2211;</mml:mo>
    <mml:mrow>
     <mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn>
    </mml:mrow>
    <mml:mi>m</mml:mi>
   </mml:munderover>
   <mml:mrow>
    <mml:msubsup>
     <mml:mi>&#x0394;</mml:mi>
     <mml:mi>i</mml:mi>
     <mml:mi>k</mml:mi>
    </mml:msubsup>
    <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo>
   </mml:mrow>
  </mml:mstyle>
 </mml:mrow>
</mml:math>
</disp-formula>
</p>
<p>where</p>
<p><disp-formula id="E6"><label>[6]</label><mml:math id="M6" display='block'>
 <mml:mrow>
  <mml:msubsup>
   <mml:mi>&#x0394;</mml:mi>
   <mml:mi>i</mml:mi>
   <mml:mi>k</mml:mi>
  </mml:msubsup>
  <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo> <mml:mtable columnalign='left'>
   <mml:mtr>
    <mml:mtd>
     <mml:mi>&#x03C8;</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msup>
      <mml:mi>S</mml:mi>
      <mml:mi>k</mml:mi>
     </mml:msup>
     <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo><mml:mo>/</mml:mo><mml:mo>&#x007C;</mml:mo><mml:msup>
      <mml:mi>S</mml:mi>
      <mml:mi>k</mml:mi>
     </mml:msup>
     <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x007C;</mml:mo><mml:mtext>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</mml:mtext><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mtext>&nbsp;&nbsp;</mml:mtext><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup>
      <mml:mi>S</mml:mi>
      <mml:mi>k</mml:mi>
     </mml:msup>
     <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo>
    </mml:mtd>
   </mml:mtr>
   <mml:mtr>
    <mml:mtd>
     <mml:mn>0</mml:mn><mml:mtext>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</mml:mtext><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi>
    </mml:mtd>
   </mml:mtr>
  </mml:mtable>
   </mml:mrow>
 </mml:mrow>
</mml:math>
</disp-formula>
</p>
<p>The value 0 = <italic>&rho;</italic> = 1 is decay constant used to simulate the evaporation of the pheromone, <italic>S<sup>k</sup></italic>(<italic>t</italic>) is the feature subset found by ant <italic>k</italic> at iteration <italic>t</italic>, and |<italic>S<sup>k</sup></italic>(<italic>t</italic>)| is its length. The pheromone is updated according to both the measure of the "goodness" of the ant's feature subset (<italic>&psi;</italic>) and the size of the subset itself. By this definition, all ants can update the pheromone.</p>
<p>The overall process of ACO feature selection can be seen in <xref ref-type="fig" rid="g2">Figure 2.</xref> The process begins by generating a number of ants, m, which are then placed randomly on the graph i.e. each ant starts with one random feature. Alternatively, the number of ants to place on the graph may be set equal to the number of features within the data; each ant starts path construction at a different feature. From these initial positions, they traverse edges probabilistically until a traversal stopping criterion is satisfied. The resulting subsets are gathered and then evaluated. If an optimal subset has been found or the algorithm has executed a certain number of times, then the process halts and outputs the best feature subset encountered. If none of these conditions hold, then the pheromone is updated, a new set of ants are created and the process iterates once more.</p>
<p>
 <fig id="g2">
					<label>Figure 2</label>
					<caption>
						<title>ACO-based feature selection algorithm.</title>						
					</caption>
					<graphic xlink:href="JCSB-02-186-g002.tif"/>
				</fig>
				</p>
				<p>The main steps of proposed feature selection algorithm are as follows:</p>
				<list id="L1" list-type="order">
				<list-item>
					<p>Generation of ants and pheromone initialization.</p>
				
				<p>&bull; Determine the population of ants.</p>
				<p>&bull; Set the intensity of pheromone trial associated with any feature.</p>
				<p>&bull; Determine the maximum of allowed iterations.</p>
				</list-item>				
				<list-item>
					<p>Ant foraging and evaluation.</p>
					<p>&bull; Any ant randomly is assigned to one feature and it should visit all features and build solutions completely.</p>
				<p>&bull; The evaluation criterion is mean square error (MSE) of the classifier. If an ant is not able to decrease the MSE of the classifier in five successive steps, it will finish its work and exit.</p>
				</list-item>				
				<list-item>
					<p>Evaluation of the selected subsets.</p>
					<p>&bull; Sort selected subsets according to classifier performance and their length. Then, select the best subset.</p>
				</list-item>				
				<list-item>
					<p>Check the stop criterion.</p>
					<p>&bull; Exit, if the number of iterations is more than the maximum allowed iteration, otherwise continue.</p>
				</list-item>				
				<list-item>
				<p>Pheromone updating.</p>
				<p>&bull; Decrease pheromone concentrations of nodes then, all ants deposit the quantity of pheromone on graph. Finally, allow the best ant to deposit additional pheromone on nodes.</p>
				</list-item>				
				<list-item>
				<p>Generation of new ants.</p>
				<p>&bull; In this step previous ants are removed and new ants are generated.</p>
				</list-item>				
				<list-item>
				<p>Go to 2 and continue.</p>
				</list-item>
				</list>
				<p>The time complexity of proposed algorithm is <italic>O</italic>(<italic>Imn</italic>), where <italic>I</italic> is the number of iterations, <italic>m</italic> the number of ants, and <italic>n</italic> the number of original features. This can be seen from <xref ref-type="fig" rid="g2">Figure 2.</xref> In the worst case, each ant selects all the features. As the heuristic is evaluated after each feature is added to the candidate subset, this will result in <italic>n</italic> evaluations per ant. After the first iteration in this algorithm, mn evaluations will have been performed. After <italic>I</italic> iterations, the heuristic will be evaluated Imn times.</p>
		</sec>	 
		 <sec>
			     <title>Particle Swarm Optimization</title>
				 <p>Particle swarm optimization is an evolutionary computation technique developed by Kennedy and Eberhart (<xref ref-type="bibr" rid="r19">Kennedy &amp; Eberhart, 2001</xref>). The original intent was to graphically simulate the graceful but unpredictable movements of a flock of birds. Initial simulations were modified to form the original version of PSO. Later, Shi introduced inertia weight into the particle swarm optimizer to produce the standard PSO (<xref ref-type="bibr" rid="r19">Kennedy &amp; Eberhart, 2001</xref>).</p>
				 <p>PSO is initialized with a population of random solutions, called "particles". Each particle is treated as a point in an n-dimensional space. The ith particle is represented as <italic>X<sub>i</sub></italic> = (<italic>x<sub>i1</sub></italic>, <italic>x<sub>i2</sub></italic>, ... , <italic>x<sub>in</sub></italic>). The best previous position (pbest, the position giving the best fitness value) of any particle is recorded and represented as <italic>P<sub>i</sub></italic> = (<italic>p<sub>i1</sub></italic>, <italic>p<sub>i2</sub></italic>, ... , <italic>p<sub>in</sub></italic>). The index of the best particle among all the particles in the population is represented by the symbol "gbest". The rate of the position change (velocity) for particle <italic>i</italic> is represented as <italic>V<sub>i</sub></italic> = (<italic>v<sub>i1</sub></italic>, <italic>v<sub>i2</sub></italic>, ... , <italic>v<sub>in</sub></italic>). The particles are manipulated according to the following equation:</p>
				 <p><disp-formula id="E7"><label>[7]</label><mml:math id="M7" display='block'>
 <mml:mrow>
  <mml:msub>
   <mml:mi>V</mml:mi>
   <mml:mi>i</mml:mi>
  </mml:msub>
  <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mi>w</mml:mi><mml:mo>.</mml:mo><mml:msub>
   <mml:mi>V</mml:mi>
   <mml:mi>i</mml:mi>
  </mml:msub>
  <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:msub>
   <mml:mi>c</mml:mi>
   <mml:mn>1</mml:mn>
  </mml:msub>
  <mml:mo>.</mml:mo><mml:msub>
   <mml:mi>r</mml:mi>
   <mml:mn>1</mml:mn>
  </mml:msub>
  <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>.</mml:mo><mml:mo stretchy='false'>[</mml:mo><mml:msub>
   <mml:mi>P</mml:mi>
   <mml:mi>i</mml:mi>
  </mml:msub>
  <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>-</mml:mo><mml:msub>
   <mml:mi>X</mml:mi>
   <mml:mi>i</mml:mi>
  </mml:msub>
  <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>]</mml:mo><mml:mo>+</mml:mo><mml:msub>
   <mml:mi>c</mml:mi>
   <mml:mn>2</mml:mn>
  </mml:msub>
  <mml:mo>.</mml:mo><mml:msub>
   <mml:mi>r</mml:mi>
   <mml:mn>2</mml:mn>
  </mml:msub>
  <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>.</mml:mo><mml:mo stretchy='false'>[</mml:mo><mml:msub>
   <mml:mi>P</mml:mi>
   <mml:mi>g</mml:mi>
  </mml:msub>
  <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>-</mml:mo><mml:msub>
   <mml:mi>X</mml:mi>
   <mml:mi>i</mml:mi>
  </mml:msub>
  <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>]</mml:mo>
 </mml:mrow>
</mml:math>
</disp-formula>
</p>
<p>where</p>
<p><disp-formula id="E8"><label>[8]</label><mml:math id="M8" display='block'>
 <mml:mrow>
  <mml:msub>
   <mml:mi>X</mml:mi>
   <mml:mi>i</mml:mi>
  </mml:msub>
  <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub>
   <mml:mi>X</mml:mi>
   <mml:mi>i</mml:mi>
  </mml:msub>
  <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:msub>
   <mml:mi>V</mml:mi>
   <mml:mi>i</mml:mi>
  </mml:msub>
  <mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo>
 </mml:mrow>
</mml:math>
</disp-formula>
</p>
<p>Where <italic>w</italic> is the inertia weight, it is a positive linear function of time changing according to the generation iteration. Suitable selection of the inertia weight provides a balance between global and local exploration and results in less iteration on average to find a sufficiently optimal solution. The acceleration constants <italic>c1</italic> and <italic>c2</italic> in equation (7) represent the weighting of the stochastic acceleration terms that pull each particle toward <italic>pbest</italic> and <italic>gbest</italic> positions. Low values allow particles to roam far from target regions before being tugged back, while high values result in abrupt movement toward, or past, target regions. <italic>r<sub>1d</sub>(t)</italic> and <italic>r<sub>2d</sub>(t)</italic> &tilde; <italic>U</italic>(0,1) are random values in the range [0,1], sampled from a uniform distribution.</p>
<p>Particles' velocities on each dimension are limited to a maximum velocity, <italic>V<sub>max</sub></italic>. It determines how large steps through the solution space each particle is allowed to take. If <italic>V<sub>max</sub></italic> is too small, particles may not explore sufficiently beyond locally good regions. They could become trapped in local optima. On the other hand, if <italic>V<sub>max</sub></italic> is too high particles might fly past good solutions.</p>
<p>The first part of equation (7) provides the "flying particles" with a degree of memory capability allowing the exploration of new search space areas. The second part is the "cognition" part, which represents the private thinking of the particle itself. The third part is the "social" part, which represents the collaboration among the particles. Equation (7) is used to calculate the particle&rsquo;s new velocity according to its previous velocity and the distances of its current position from its own best experience (position) and the group's best experience. Then the particle flies toward a new position according to equation (8). The performance of each particle is measured according to a pre-defined fitness function.</p>			
			 <sec>
			     <title>PSO for Feature Selection</title>
				 <p>The idea of PSO can be used for the optimal feature selection problem. Consider a large feature space full of feature subsets. Each feature subset can be seen as a point or position in such a space. If there are <italic>n</italic> total features, then there will be 2<sup>n</sup> kinds of subsets, different from each other in the length and features contained in each subset. The optimal position is the subset with the least length and the highest classification accuracy. A particle swarm is put into this feature space, each particle takes one position. The particles fly in this space, their goal is to fly to the best position. After a while, they change their position, communicate with each other, and search around the local best and global best position. Eventually, they should converge on good, possibly optimal, positions. It is the exploration ability of particle swarms that equips them for performing feature selection and discovering optimal subsets.</p>
			 </sec>
		 </sec>
		 <sec>
		     <title>Experimental Results</title>
			 <p>In this section, we report and discuss computational experiments and compare ACO feature selection algorithm with PSO-based approach. The quality of a candidate solution is computed by the naive Bayes classifier which described in section 2. Finally the classifier performance and the length of selected feature subset are considered for evaluating the proposed algorithm.</p>
			 <p>A series of experiments was conducted to show the utility of proposed feature selection algorithm. All experiments have been run on a machine with 3.0GHz CPU and 1024 MB of RAM. We implement ACO algorithm and PSObased algorithm in Java and Weka 3.5.5. The operating system was Windows XP Professional. For experimental studies we have considered Postsynaptic dataset. The following sections describe Postsynaptic dataset and implementation results.</p>
			 <p>
			 <fig id="g3">
			 <label>Figure 3</label>
			 <caption>
			 <title>Main elements involved in pre-synaptic and postsynaptic activity.</title>						
			 </caption>
			 <graphic xlink:href="JCSB-02-186-g003.tif"/>
			 </fig>
			 </p>			
			 <sec>
			     <title>Postsynaptic Dataset</title>
				 <p>This section presents Postsynaptic dataset used in the present work for feature selection. A synapse is a connection between two neurons: pre-synaptic and post-synaptic. The first is usually the sender of some signals such as the release of chemicals, while the second is the receiver. A post-synaptic receptor is a sensor on the surface of a neuron. It captures messenger molecules from the nervous system, neurotransmitters, and thereby functions in transmitting information from one neuron to another (<xref ref-type="bibr" rid="r26">Pappa et al., 2005</xref>).</p>
				 <p>The main elements found in synapses are shown in <xref ref-type="fig" rid="g3">Figure 3.</xref> The cells are held together by cell adhesion molecules (1). In the cell where the signal is coming from (the pre-synaptic cell) neurotransmitters are stored in bags called synaptic vesicles. When signals are to be transmitted from the pre-synaptic cell to the post-synaptic cell, synaptic vesicles fuse with the pre-synaptic membrane and release their contents into the synaptic cleft between the cells. The transmitters then diffuse within the cleft, and some of them meet a post-synaptic receptor (2), which recognizes them as a signal. This activates the receptor, which then transmits the signal on to other signalling components such as voltage-gated ion channels (3), protein kinases (4) and phosphatases (5). To ensure that the signal has terminated, transporters (6) remove neurotransmitters from the cleft. Within the post-synaptic cell, the signalling apparatus is organized by various scaffolding proteins (7).</p>
				 <p>Postsynaptic dataset has been recently created and mined in (<xref ref-type="bibr" rid="r3">Basiri et al., 2008</xref>; <xref ref-type="bibr" rid="r5">Correa et al., 2006</xref>; <xref ref-type="bibr" rid="r26">Pappa et al., 2005</xref>). The dataset contains 4303 records of proteins. These proteins belong to either positive or negative classes. Proteins that belong to the positive class have post-synaptic activity while negative ones don't show such activity. From the 4303 proteins on the dataset, 260 belong to the positive class and 4043 to the negative class. This dataset has many features which makes the feature selection task challenging. More precisely, each protein has 443 PROSITE patterns, or features. PROSITE is a database of protein families and domains. It is based on the observation that, while there are a huge number of different proteins, most of them can be grouped, on the basis of similarities in their sequences, into a limited number of families (a protein consists of a
sequence of amino acids). PROSITE patterns are small regions within a protein that present a high sequence similarity when compared to other proteins. In our dataset the absence of a given PROSITE pattern is indicated by a value of 0 for the feature corresponding to that PROSITE pattern which its presence is indicated by a value of 1 for that same feature (<xref ref-type="bibr" rid="r5">Correa et al., 2006</xref>).</p>
			 </sec>	
			 <sec sec-type="methods">
		    <title>Experimental Methodology</title>
			<p>The computational experiments involved a ten-fold crossvalidation method (<xref ref-type="bibr" rid="r32">Witten &amp; Frank, 2005</xref>). First, the 4303 records in the Postsynaptic dataset were divided into 10 almost equally sized folds. There are three folds containing 431 records each one and seven folds containing 430 records each one. The folds were randomly generated but under the following regulation. The proportion of positive and negative classes in every single fold must be similar to the one found in the original dataset containing all the 4303 records. This is known as stratified cross-validation. Each of the 10 folds is used once as test set and the remaining of the dataset is used as training set. Out of the 9 folds in the training set, one is reserved to be used as a validation set.</p>
			<p>In each of the 10 iterations of the cross-validation procedure, the predictive accuracy of the classification is assessed by 3 different methods:</p>
			<list id="l8" list-type="bullet">
			<list-item>
				<p><bold>Using all the 443 original features:</bold> all possible features are used by the nearest neighbor classifier and the naive Bayes classifier.</p>
			</list-item>
			<list-item>
				<p><bold>Standard binary PSO algorithm:</bold> only the features selected by the best particle found by the binary PSO algorithm are used by the nearest neighbor classifier and the naive Bayes classifier.</p>
			</list-item>
			<list-item>
				<p><bold>Proposed ACO algorithm:</bold> only the features selected by the best ant found by the ACO algorithm are used by the nearest neighbor classifier and the naive Bayes classifier.</p>
			</list-item>
			</list>
			<p>As the standard binary PSO and the ACO algorithm are stochastic algorithms, 20 independent runs for each algorithm were performed for every iteration of the cross-validation procedure. The obtained results, averaged over 20  runs, are reported in <xref ref-type="table" rid="t2">Table 2.</xref> The average number of features selected by the feature selection algorithms has always been rounded to the nearest integer.</p>
			<p>Various values were tested for the parameters of ACO algorithm. The results show that the highest performance is achieved by setting the parameters to values shown in <xref ref-type="table" rid="t1">Table 1.</xref></p>
			<p>Where w is inertia weight and <italic>c</italic>1 and <italic>c</italic>2 are acceleration constants of standard binary PSO algorithm. The choice of the value of this parameter was based on the work presented in (<xref ref-type="bibr" rid="r29">Shi, &amp; Eberhart, 1998</xref>). For ant colony optimization, parameter values were empirically determined in our preliminary experiments for leading to better convergence; but we make no claim that these are optimal values. Parameter optimization is a topic for future research.</p>
			</sec>
			<sec>
			    <title>Performance Measure</title>
				<p>The measurement of the predictive accuracy rate of a model should be a reliable estimate of how well that model classifies the test examples (unseen during the training phase) on the target problem. In data mining, typically, the following equation is used to assess the accuracy rate of a classifier:
				 <disp-formula id="E9"><label>[9]</label><mml:math id="M9" display='block'>
 <mml:semantics>
  <mml:mrow>
   <mml:mtext>Standard&nbsp;accuracy&nbsp;rate</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x2009;</mml:mtext><mml:mfrac>
    <mml:mrow>
     <mml:mtext>TP+TN</mml:mtext>
    </mml:mrow>
    <mml:mrow>
     <mml:mtext>TP+FP+FN+TN</mml:mtext>
    </mml:mrow>
   </mml:mfrac>
   
  </mml:mrow>
 <mml:annotation encoding='MathType-MTEF'>
 </mml:annotation>
 </mml:semantics>
</mml:math>
</disp-formula>
</p>
<p>where <italic>TP</italic> (true positives) is the number of records correctly classified as positive class and <italic>FP</italic> (false positives) is the number of records incorrectly classified as positive class. <italic>TN</italic> (true negatives) is the number of records correctly classified as negative class and <italic>FN</italic> (false negatives) is the number of records incorrectly classified as negative class.</p>
<p>Nevertheless, if the class distribution is highly unbalanced, which is the case with the Postsynaptic dataset, equation (9) is an ineffective way of measuring the accuracy rate of a model. For instance, on a dataset in which 10% of the examples belong to the positive class and 90% to the negative class, it would be easy to maximize equation (9) by simply predicting always the majority class. Therefore, on our experiments we use a more demanding measurement for the accuracy rate of a classification model. It has also been used before in (<xref ref-type="bibr" rid="r3">Basiri et al., 2008</xref>; <xref ref-type="bibr" rid="r26">Pappa et al., 2005</xref>). This measurement is given by the equation:<disp-formula id="E10"><label>[10]</label><mml:math id="M10" display='block'>
 <mml:semantics>
  <mml:mrow>
   <mml:mtext>Predictive&nbsp;accuracy&nbsp;rate&nbsp;=&nbsp;TPR&#x00D7;TNR</mml:mtext>
  </mml:mrow>
 <mml:annotation encoding='MathType-MTEF'>
 </mml:annotation>
 </mml:semantics>
</mml:math></disp-formula></p>
<p>where TPR and TNR are defined as follows:</p>
<p><disp-formula id="E11"><label>[11]</label><mml:math id="M11" display='block'>
 <mml:mrow>
  <mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac>
   <mml:mrow>
    <mml:mi>T</mml:mi><mml:mi>P</mml:mi>
   </mml:mrow>
   <mml:mrow>
    <mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi>
   </mml:mrow>
  </mml:mfrac>  
 </mml:mrow>
</mml:math>
</disp-formula>
</p>
<p><disp-formula id="E12"><label>[12]</label><mml:math id="M12" display='block'>
 <mml:mrow>
  <mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac>
   <mml:mrow>
    <mml:mi>T</mml:mi><mml:mi>N</mml:mi>
   </mml:mrow>
   <mml:mrow>
    <mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi>
   </mml:mrow>
  </mml:mfrac>  
 </mml:mrow>
</mml:math>
</disp-formula>
</p>
			</sec>
			<sec>
		     <title>Results</title>
			 <p><xref ref-type="table" rid="t2">Table 2.</xref> gives the optimal selected features for each method. As discussed earlier the experiments involved 200 runs of ACO and standard binary PSO, 10 cross-validation folds times 20 runs with different random seeds. Presumably, those 200 runs selected different subsets of features. So, the features which have been listed in <xref ref-type="table" rid="t2">Table 2.</xref> are the ones most often selected by ACO and standard binary PSO across all the 20 runs. Both ACO-based and PSO-based methods significantly reduce the number of original features, however; ACO-based method chooses fewer features.</p>
			 <p>Another trend observed in the results was found at each run of the ACO algorithm. We recorded the best ant by the ACO algorithm on each of the 20 runs and for every one of the 10 folds. We then computed the frequency of the features selected on the 10 &chi; 20 = 200 best ants found by the ACO algorithm. The following 3 features have been selected, by ACO algorithm, in more than 80% of its respective 200 best ants found: <italic>F<sub>342</sub>, F<sub>352</sub> and F<sub>353</sub></italic>. The names the PROSITE patterns that correspond to these features are shown in <xref ref-type="table" rid="t3">Table 3.</xref></p>
			 <p>The information was obtained from the web site of the European Bioinformatics Institute, UniProtKB/Swiss-Prot (<ext-link ext-link-type="uri" xlink:href="www.ebi.ac.uk/swissprot">http://www.ebi.ac.uk/swissprot/</ext-link>).</p>
			 <p>Also, the results of both algorithms for all of the 10 folds are summarized in <xref ref-type="table" rid="t4">Table 4.</xref> and <xref ref-type="table" rid="t5">Table 5.</xref> The classification quality and feature subset length are two criteria which are considered to assess the performance of algorithms. Comparing these criteria, we noted that ACO and standard binary PSO algorithms did very better than the Baseline algorithm (using all features). Furthermore, for all of the 10 folds the ACO algorithm selected a smaller subset of features than the standard binary PSO algorithm.</p>
			 <p>As we can see in <xref ref-type="table" rid="t5">Table 5.</xref>, the average number of selected features for standard binary PSO algorithm was equal to 10.9 with the average predictive accuracy of 0.79 and the average number of selected features for ACO algorithm was equal to 4.2 with the average predictive accuracy of 0.85. Furthermore, in (<xref ref-type="bibr" rid="r5">Correa et al., 2006</xref>) a new discrete PSO algorithm, called DPSO, has been introduced for feature selection. DPSO has been applied to Postsynaptic dataset and the average number of features selected by that was 12.70 with the average predictive accuracy of 0.74. Comparison of these three algorithms shows that ACO tends to select a smaller subset of features than the standard binary PSO algorithm and DPSO. Also, the average predictive accuracy of ACO is higher than that of the standard binary PSO algorithm and DPSO. Predictive accuracy and number of selected features for ACO and standard binary PSO algorithm are shown in <xref ref-type="fig" rid="g4">Figure 4.</xref> and <xref ref-type="fig" rid="g5">Figure 5.</xref>.</p>
			 <p>
			 <fig id="g4">
			 <label>Figure 4</label>
			 <caption>
			 <title>(a) Accuracy rate of feature subsets obtained using three methods. (b) Number selected features. (Using nearest
neighbor classifier)</title>						
			 </caption>
			 <graphic xlink:href="JCSB-02-186-g004.tif"/>
			 </fig>
			 </p>
			 <p>
			 <fig id="g5">
			 <label>Figure 5</label>
			 <caption>
			 <title>(a) Accuracy rate of feature subsets obtained using three methods. (b) Number selected features. (Using naive Bayes classifier)</title>
			 </caption>
			 <graphic xlink:href="JCSB-02-186-g005.tif"/>
			 </fig>
			 </p>
		 </sec>
		 
		 <sec>
		     <title>Discussion</title>
			 <p>Experimental results show that the use of unnecessary features hurt classification accuracy and FS is used to reduce redundancy in the information provided by the selected features. Using only a small subset of selected features, the ACO and the PSO algorithms obtained better classification accuracy than the baseline algorithm using all features.</p>
			 <p>ACO shares many similarities with evolutionary computation (EC) techniques in general and GAs in particular. These techniques begin with a group of a randomly generated population and utilize a fitness value to evaluate the population. They all update the population and search for the optimum with random techniques.</p>
			 <p>Both ACO and PSO are stochastic population-based search approaches that depend on information sharing among their population members to enhance their search processes using a combination of deterministic and probabilistic rules. They are efficient, adaptive and robust search processes, producing near optimal solutions, and have a large degree of implicit parallelism. The main difference between the ACO compared to PSO, is that ACO does not have PSO operators such as inertia weight and acceleration constants. Ants update themselves with the pheromone update rule; they also have a memory that is important to the algorithm.</p>
			 <p>Compared to PSO, the ACO has a much more intelligent background and can be implemented more easily. The computation time used in ACO is less than in PSO. The parameters used in ACO are also fewer. However, if the proper parameter values are set, the results can easily be optimized. The decision on the parameters of the ant colony affects the exploration&ndash;exploitation tradeoff and is highly dependent on the form of the objective function. Successful feature selection was obtained even using conservative values for the ACO basic parameters.</p>
		  </sec>
		  </sec>
		  <sec>
		     <title>Conclusion</title>
			 <p>Experimental results show that the use of unnecessary features decrease classifiers' performance and hurt classification accuracy and feature selection is used to reduces redundancy in the information provided by the selected features. Using only a small subset of selected features, the binary PSO and the ACO algorithms obtained better predictive accuracy than the Baseline algorithm using all features. ACO has the ability to converge quickly; it has a strong search capability in the problem space and can efficiently find minimal feature subset. Experimental results demonstrate competitive performance.</p>
			 <p>The ACO clearly enhances computational efficiency of the classifier by selecting fewer features than the standard binary PSO algorithm. Therefore, when the difference in predictive accuracy is insignificant, ACO is still preferable. In the proposed ACO algorithm, the classifier performance and the length of selected feature subset are adopted as heuristic information. So, we can select the optimal feature subset without the prior knowledge of features.</p>
			 <p>Also as we expected, computational results show the clear difference in performance between nearest neighbor and naive Bayes classifiers. To summarize, the naive Bayes classification method involves a classification step in which the various <italic>P</italic>(<italic>v<sub>i</sub></italic>) and <italic>P</italic>(<italic>a<sub>i</sub></italic> | <italic>v<sub>i</sub></italic>) terms are estimated, based on their frequencies over the training data. The set of these estimates corresponds to the learned hypothesis. This hypothesis is then used to classify each new instance by applying the rule in equation (3). Whenever the naive Bayes assumption of conditional independence is satisfied, this naive Bayes classification <italic>v<sub>NB</sub></italic> is identical to the maximum a posterior classification. The naive Bayes approach outperformed the nearest neighbor approach in all experiments.</p>
		 </sec>
		</body>
	<back>
	<ack>
			<p>The authors wish to thank the Office of Graduate studies of the University of Isfahan for their support and also wish to offer their special thanks to Dr. Alex A. Freitas for providing Postsynaptic dataset.</p>
	</ack>
	<ref-list>
	   <title>References</title>
	         <ref id="r1">
			 <citation citation-type="confproc">
			         <person-group>
					 <name>
					      <surname>Aghdam</surname>
						  <given-names>MH</given-names>
					</name>
					<name>
					      <surname>Ghasem-</surname>
						  <given-names>AghaeeN</given-names>
					</name>
					<name>
					      <surname>Basiri</surname>
						  <given-names>ME</given-names>
					</name>
					</person-group>
					<year>2008</year>
					<article-title>Application of Ant Colony Optimization for Feature Selection in Text Categorization</article-title>
					<conf-name>Proceedings of the IEEE Congress on Evolutionary Computation (CEC 2008)</conf-name>
					<fpage>pp 2872</fpage>
					<lpage>2878</lpage>
			</citation>
			</ref>
			<ref id="r2">
			 <citation citation-type="book">
			         <person-group>
					 <name>
					      <surname>Alberts</surname>
						  <given-names>B</given-names>
					</name>
					<name>
					      <surname>Bruce</surname>
						  <given-names>A</given-names>
					</name>
					<name>
					      <surname>Johnson</surname>
						  <given-names>A</given-names>
					</name>
					<name>
					      <surname>Lewis</surname>
						  <given-names>J</given-names>
					</name>
					<name>
					      <surname>Raff</surname>
						  <given-names>M</given-names>
					</name><etal/>
					</person-group>
					<year>2002</year>
					<article-title>The molecular biology of the cell</article-title>
					<edition>4th</edition>
					<publisher-name>Garland Press</publisher-name>
			</citation>
			</ref>
			<ref id="r3">
			 <citation citation-type="confproc">
			         <person-group>
					 <name>
					      <surname>Basiri</surname>
						  <given-names>ME</given-names>
					</name>
					<name>
					      <surname>Ghasem-</surname>
						  <given-names>AghaeeN</given-names>
					</name>
					<name>
					      <surname>Aghdam</surname>
						  <given-names>MH</given-names>
					</name>
					</person-group>
					<year>2008</year>
					<article-title>Using ant colony optimization-based selected features for predicting post-synaptic activity in proteins</article-title>
					<conf-name>Proceedings of 6th European Conference on Evolutionary Computation, Machine Learning and Data Mining in Bioinformatics</conf-name>
					<conf-loc>LNCS 4973</conf-loc>
					<fpage>pp 12</fpage>
					<lpage>23</lpage>
			</citation>
			</ref>
			<ref id="r4">
			 <citation citation-type="book">
			         <person-group>
					 <name>
					      <surname>Bonabeau</surname>
						  <given-names>E</given-names>
					</name>
					<name>
					      <surname>Dorigo</surname>
						  <given-names>M</given-names>
					</name>
					<name>
					      <surname>Theraulaz</surname>
						  <given-names>G</given-names>
					</name>					
					</person-group>
					<year>1999</year>
					<article-title>Swarm Intelligence: From Natural to Artificial Systems</article-title>								
					<publisher-name>Oxford University Press</publisher-name>
					<publisher-loc>New York</publisher-loc>					
			</citation>
			</ref>
			<ref id="r5">
			 <citation citation-type="confproc">
			         <person-group>
					 <name>
					      <surname>Correa</surname>
						  <given-names>ES</given-names>
					</name>
					<name>
					      <surname>Freitas</surname>
						  <given-names>AA</given-names>
					</name>
					<name>
					      <surname>Johnson</surname>
						  <given-names>CG</given-names>
					</name>
					</person-group>
					<year>2006</year>
					<article-title>A new discrete particle swarm algorithm applied to attribute selection in a bioinformatics dataset</article-title>
					<conf-name>Proceedings of the Genetic and Evolutionary Computation Conference</conf-name>					
					<fpage>pp 35</fpage>
					<lpage>42</lpage>
			</citation>
			</ref>
			<ref id="r6">
			 <citation citation-type="journal">
			         <person-group>
					 <name>
					      <surname>Dash</surname>
						  <given-names>M</given-names>
					</name>
					<name>
					      <surname>Liu</surname>
						  <given-names>H</given-names>
					</name>					
					</person-group>
					<year>1997</year>
					<article-title>Feature Selection for Classification</article-title>
					<source>Intelligent Data Analysis</source>
					<volume>3</volume>					
					<fpage>131</fpage>
					<lpage>156</lpage>
			</citation>
			</ref>
			<ref id="r7">
			 <citation citation-type="journal">
			         <person-group>
					 <name>
					      <surname>Domeniconi</surname>
						  <given-names>C</given-names>
					</name>
					 <name>
					      <surname>Peng</surname>
						  <given-names>J</given-names>
					</name>
					<name>
					      <surname>Gunopulos</surname>
						  <given-names>D</given-names>
					</name>					
					</person-group>
					<year>2002</year>
					<article-title>Locally adaptive metric nearest-neighbor classification</article-title>
					<source>IEEE Transaction Pattern Analysis and Machine Intelligence</source>
					<volume>24</volume>					
					<fpage>1281</fpage>
					<lpage>1285</lpage>
			</citation>
			</ref>
			<ref id="r8">
			 <citation citation-type="confproc">
			         <person-group>
					 <name>
					      <surname>Dorigo</surname>
						  <given-names>M</given-names>
					</name>					
					</person-group>
					<year>1992</year>
					<article-title>Optimization, learning and natural algorithms</article-title>
					<conf-name>PhD thesis</conf-name>
					<conf-loc>Dipartimento di Elettronica</conf-loc>
					<publisher-name>Politecnico di Milano</publisher-name>
					<publisher-loc>Italy</publisher-loc>										
			</citation>
			</ref>
			<ref id="r9">
			 <citation citation-type="confproc">
			         <person-group>
					 <name>
					      <surname>Dorigo</surname>
						  <given-names>M</given-names>
					</name>
					<name>
					      <surname>Caro</surname>
						  <given-names>GD</given-names>
					</name>					
					</person-group>
					<year>1999</year>
					<article-title>Ant Colony Optimization: A New Meta-heuristic</article-title>
					<conf-date>1999</conf-date>
					<conf-name>Proceedings of IEEE Congress on Evolutionary Computing</conf-name>															
			</citation>
			</ref>
			<ref id="r10">
			 <citation citation-type="journal">
			         <person-group>
					 <name>
					      <surname>Dorigo</surname>
						  <given-names>M</given-names>
					</name>
					 <name>
					      <surname>Maniezzo</surname>
						  <given-names>V</given-names>
					</name>
					<name>
					      <surname>Colorni</surname>
						  <given-names>A</given-names>
					</name>					
					</person-group>
					<year>1996</year>
					<article-title>The Ant System: Optimization by a colony of cooperating agents</article-title>
					<source>IEEE Transaction on Systems, Man, and Cybernetics- Part B</source>
					<volume>26</volume>					
					<fpage>29</fpage>
					<lpage>41</lpage>
			</citation>
			</ref>
			<ref id="r11">
			 <citation citation-type="book">
			         <person-group>
					 <name>
					      <surname>Engelbrecht</surname>
						  <given-names>AP</given-names>
					</name>										
					</person-group>
					<year>2005</year>
					<article-title>Fundamentals of Computational Swarm Intelligence</article-title>										
					<publisher-name>Wiley</publisher-name>
					<publisher-loc>London</publisher-loc>					
			</citation>
			</ref>
			<ref id="r12">
			 <citation citation-type="book">
			         <person-group>
					 <name>
					      <surname>Feller</surname>
						  <given-names>W</given-names>
					</name>										
					</person-group>
					<year>1971</year>
					<article-title>An introduction to probability theory and its applications</article-title>										
					<publisher-name>John Wiley and Sons, Inc, New York</publisher-name>
					<publisher-loc>NY, USA</publisher-loc>					
			</citation>
			</ref>
			<ref id="r13">
			 <citation citation-type="journal">
			         <person-group>
					 <name>
					      <surname>Forman</surname>
						  <given-names>G</given-names>
					</name>										
					</person-group>
					<year>2003</year>
					<article-title>An extensive empirical study of feature selection metrics for text classification</article-title>										
					<source>Journal of Machine Learning Research</source>
					<fpage>1289</fpage>
					<lpage>1305</lpage>					
			</citation>
			</ref>
			<ref id="r14">
			 <citation citation-type="journal">
			         <person-group>
					 <name>
					      <surname>Freitas</surname>
						  <given-names>AA</given-names>
					</name>
					<name>
					      <surname>de Carvalho</surname>
						  <given-names>ACPLF</given-names>
					</name>										
					</person-group>
					<year>2007</year>
					<article-title>A tutorial on hierarchical classification with applications in bioinformatics</article-title>										
					<source>Research and Trends in Data Mining Technologies and Applications</source>
					<volume>99</volume>
					<fpage>175</fpage>
					<lpage>208</lpage>					
			</citation>
			</ref>
			<ref id="r15">
			 <citation citation-type="journal">
			         <person-group>
					 <name>
					      <surname>Guyon</surname>
						  <given-names>I</given-names>
					</name>
					<name>
					      <surname>Elisseeff</surname>
						  <given-names>A</given-names>
					</name>										
					</person-group>
					<year>2003</year>
					<article-title>An Introduction to Variable and Feature Selection</article-title>										
					<source>Journal of Machine Learning Research</source>
					<volume>3</volume>
					<fpage>1157</fpage>
					<lpage>1182</lpage>					
			</citation>
			</ref>
			<ref id="r16">
			 <citation citation-type="book">
			         <person-group>
					 <name>
					      <surname>Han</surname>
						  <given-names>J</given-names>
					</name>
					<name>
					      <surname>Kamber</surname>
						  <given-names>M</given-names>
					</name>										
					</person-group>
					<year>2001</year>
					<article-title>Data mining: concepts and techniques</article-title>										
					<publisher-name>Morgan Kaufmann Publishers, San Francisco</publisher-name>
					<publisher-loc>CA, USA</publisher-loc>					
			</citation>
			</ref>
			<ref id="r17">
			 <citation citation-type="book">
			         <person-group>
					 <name>
					      <surname>Jensen</surname>
						  <given-names>R</given-names>
					</name>															
					</person-group>
					<year>2005</year>
					<article-title>Combining rough and fuzzy sets for feature selection</article-title>										
					<publisher-name>Ph.D. dissertation, School of Information</publisher-name>
					<publisher-loc>Edinburgh University</publisher-loc>					
			</citation>
			</ref>
			<ref id="r18">
			 <citation citation-type="confproc">
			         <person-group>
					 <name>
					      <surname>John</surname>
						  <given-names>GH</given-names>
					</name>
					<name>
					      <surname>Kohavi</surname>
						  <given-names>R</given-names>
					</name>
					<name>
					      <surname>Pfleger</surname>
						  <given-names>K</given-names>
					</name>					
					</person-group>
					<year>1994</year>
					<article-title>Irrelevant Features and the Subset Selection Problem</article-title>
					<conf-name>Proceedings of the 11th International Conference on Machine Learning ICML</conf-name>
					<volume>94</volume>
					<fpage>121</fpage>
					<lpage>129</lpage>										
			</citation>
			</ref>
			<ref id="r19">
			 <citation citation-type="book">
			         <person-group>
					 <name>
					      <surname>Kennedy</surname>
						  <given-names>J</given-names>
					</name>
					<name>
					      <surname>Eberhart</surname>
						  <given-names>RC</given-names>
					</name>															
					</person-group>
					<year>2001</year>
					<article-title>Swarm Intelligence</article-title>										
					<publisher-name>Morgan Kaufmann Publishers Inc</publisher-name>
					<publisher-loc>San Francisco</publisher-loc>					
			</citation>
			</ref>
			<ref id="r20">
			 <citation citation-type="journal">
			         <person-group>
					 <name>
					      <surname>Kohavi</surname>
						  <given-names>R</given-names>
					</name>
					<name>
					      <surname>John</surname>
						  <given-names>GH</given-names>
					</name>										
					</person-group>
					<year>1997</year>
					<article-title>Wrappers for feature subset selection</article-title>										
					<source>Artificial Intelligence</source>
					<volume>97</volume>
					<fpage>273</fpage>
					<lpage>324</lpage>					
			</citation>
			</ref>
			<ref id="r21">
			 <citation citation-type="confproc">
			         <person-group>
					 <name>
					      <surname>Latourrette</surname>
						  <given-names>M</given-names>
					</name>										
					</person-group>
					<year>2000</year>
					<article-title>Toward an explanatory similarity measure for nearest-neighbor classification</article-title>
					<conf-name>In: ECML '00: Proceedings of the 11th European Conference on Machine Learning</conf-name>
					<conf-loc>London UK</conf-loc>					
					<fpage>238</fpage>
					<lpage>245</lpage>										
			</citation>
			</ref>
			<ref id="r22">
			 <citation citation-type="journal">
			         <person-group>
					 <name>
					      <surname>Liu</surname>
						  <given-names>B</given-names>
					</name>
					<name>
					      <surname>Abbass</surname>
						  <given-names>HA</given-names>
					</name>
					<name>
					      <surname>McKay</surname>
						  <given-names>B</given-names>
					</name>										
					</person-group>
					<year>2004</year>
					<article-title>Classification Rule Discovery with Ant Colony Optimization</article-title>										
					<source>IEEE Computational Intelligence Bulletin</source>
					<volume>3</volume>
					<fpage>31</fpage>
					<lpage>35</lpage>					
			</citation>
			</ref>
			<ref id="r23">
			 <citation citation-type="book">
			         <person-group>
					 <name>
					      <surname>Liu</surname>
						  <given-names>H</given-names>
					</name>
					<name>
					      <surname>Motoda</surname>
						  <given-names>H</given-names>
					</name>															
					</person-group>
					<year>1998</year>
					<article-title>Feature Selection for Knowledge Discovery and Data Mining</article-title>										
					<publisher-name>Boston: Kluwer Academic Publishers</publisher-name>										
			</citation>
			</ref>
			<ref id="r24">
			 <citation citation-type="journal">
			         <person-group>
					 <name>
					      <surname>Maniezzo</surname>
						  <given-names>V</given-names>
					</name>
					<name>
					      <surname>Colorni</surname>
						  <given-names>A</given-names>
					</name>															
					</person-group>
					<year>1999</year>
					<article-title>The Ant System Applied to the Quadratic Assignment Problem</article-title>										
					<source>IEEE Transaction on Knowledge and Data Engineering</source>
					<volume>11</volume>
					<fpage>pp769</fpage>
					<lpage>778</lpage>					
			</citation>
			</ref>
			<ref id="r25">
			 <citation citation-type="book">
			         <person-group>
					 <name>
					      <surname>Mitchell</surname>
						  <given-names>T</given-names>
					</name>																				
					</person-group>
					<year>1996</year>
					<article-title>Machine Learning</article-title>										
					<publisher-name>McCraw Hill</publisher-name>										
			</citation>
			</ref>
			<ref id="r26">
			 <citation citation-type="journal">
			         <person-group>
					 <name>
					      <surname>Pappa</surname>
						  <given-names>GL</given-names>
					</name>
					<name>
					      <surname>Baines</surname>
						  <given-names>AJ</given-names>
					</name>
					<name>
					      <surname>Freitas</surname>
						  <given-names>AA</given-names>
					</name>															
					</person-group>
					<year>2005</year>
					<article-title>Predicting postsynaptic activity in proteins with data mining</article-title>										
					<source>Bioinformatics</source>
					<volume>21</volume>
					<fpage>pp19</fpage>
					<lpage>25</lpage>					
			</citation>
			</ref>
			<ref id="r27">
			 <citation citation-type="book">
			         <person-group>
					 <name>
					      <surname>Pawlak</surname>
						  <given-names>Z</given-names>
					</name>																				
					</person-group>
					<year>1991</year>
					<article-title>Rough Sets: Theoretical Aspects of Reasoning about Data</article-title>										
					<publisher-name>Kluwer Academic Publishing</publisher-name>
					<publisher-loc>Dordrecht</publisher-loc>										
			</citation>
			</ref>
			<ref id="r28">
			 <citation citation-type="journal">
			         <person-group>
					 <name>
					      <surname>Pudil</surname>
						  <given-names>P</given-names>
					</name>
					<name>
					      <surname>Novovicova</surname>
						  <given-names>J</given-names>
					</name>
					<name>
					      <surname>Kittler</surname>
						  <given-names>J</given-names>
					</name>															
					</person-group>
					<year>1994</year>
					<article-title>Floating search methods in feature selection</article-title>										
					<source>Pattern Recognition Letters</source>
					<volume>15</volume>
					<fpage>pp1119</fpage>
					<lpage>1125</lpage>					
			</citation>
			</ref>
			<ref id="r29">
			 <citation citation-type="confproc">
			         <person-group>
					 <name>
					      <surname>Shi</surname>
						  <given-names>Y</given-names>
					</name>
					<name>
					      <surname>Eberhart</surname>
						  <given-names>RC</given-names>
					</name>										
					</person-group>
					<year>1998</year>
					<article-title>Parameter selection in particle swarm optimization</article-title>
					<conf-name>Proceedings of the 7th International Conference on Evolutionary Programming</conf-name>								
					<fpage>591</fpage>
					<lpage>600</lpage>										
			</citation>
			</ref>
			<ref id="r30">
			 <citation citation-type="journal">
			         <person-group>
					 <name>
					      <surname>Siedlecki</surname>
						  <given-names>W</given-names>
					</name>
					<name>
					      <surname>Sklansky</surname>
						  <given-names>J</given-names>
					</name>																				
					</person-group>
					<year>1988</year>
					<article-title>On Automatic Feature Selection</article-title>										
					<source>International Journal of Pattern Recognition and Artificial Intelligence</source>
					<volume>2</volume>
					<fpage>pp197</fpage>
					<lpage>220</lpage>					
			</citation>
			</ref>
			<ref id="r31">
			 <citation citation-type="journal">
			         <person-group>
					 <name>
					      <surname>Wang</surname>
						  <given-names>X</given-names>
					</name>
					<name>
					      <surname>Yang</surname>
						  <given-names>J</given-names>
					</name>
					<name>
					      <surname>Teng</surname>
						  <given-names>X</given-names>
					</name>
					<name>
					      <surname>Xia</surname>
						  <given-names>W</given-names>
					</name>
					<name>
					      <surname>Jensen</surname>
						  <given-names>R</given-names>
					</name>																				
					</person-group>
					<year>2007</year>
					<article-title>Feature selection based on rough sets and particle swarm optimization</article-title>										
					<source>Pattern Recognition Letters</source>
					<volume>28</volume>
					<fpage>pp459</fpage>
					<lpage>471</lpage>					
			</citation>
			</ref>
			<ref id="r32">
			 <citation citation-type="book">
			         <person-group>
					 <name>
					      <surname>Witten</surname>
						  <given-names>IH</given-names>
					</name>
					<name>
					      <surname>Frank</surname>
						  <given-names>E</given-names>
					</name>																				
					</person-group>
					<year>2005</year>
					<article-title>Data Mining: Practical Machine Learning Tools and Techniques</article-title>										
					<publisher-name>Morgan Kaufmann</publisher-name>
					<publisher-loc>San Francisco</publisher-loc>										
			</citation>
			</ref>
			<ref id="r33">
			 <citation citation-type="journal">
			         <person-group>
					 <name>
					      <surname>Yang</surname>
						  <given-names>J</given-names>
					</name>
					<name>
					      <surname>Honavar</surname>
						  <given-names>V</given-names>
					</name>																				
					</person-group>
					<year>1998</year>
					<article-title>Feature Subset Selection Using a Genetic Algorithm</article-title>										
					<source>IEEE Intelligent Systems</source>
					<volume>13</volume>
					<fpage>pp44</fpage>
					<lpage>49</lpage>					
			</citation>
			</ref>
	</ref-list>
	</back>
	<floats-wrap>
	<table-wrap position="float" id="t1">
	<label>Table 1.</label>
  			<caption>
  				<title>Binary PSO and ACO parameter settings.</title>
  			</caption>
   <table>
      <thead>
		 <tr>
		 	<th>Method</th>
			<th>Population</th>
			<th>Iteration</th>
			<th>Initial pheromone</th>
			<th><italic>c</italic><sub>1</sub></th>
			<th><italic>c</italic><sub>2</sub></th>
			<th><italic>w</italic></th>
			<th><italic>&alpha;</italic></th>
			<th><italic>&beta;</italic></th>
			<th><italic>&rho;</italic></th>
		</tr>		 
      </thead>
      <tbody>
	     <tr>
			<td>BPSO</td>	
			<td>30</td>
			<td>50</td>
			<td>-</td>									
			<td>2</td>
			<td>2</td>
			<td>0.8</td>
			<td>-</td>
			<td>-</td>
			<td>-</td>
         </tr>        
         <tr>
            <td>ACO</td>	
			<td>30</td>
			<td>50</td>
			<td>1</td>									
			<td>-</td>
			<td>-</td>
			<td>-</td>
			<td>1</td>
			<td>0.1</td>
			<td>0.2</td>										
         </tr>
	</tbody>
	</table>
	</table-wrap>
	<table-wrap position="float" id="t2">
	<label>Table 2.</label>
  			<caption>
  				<title>Selected features of standard binary PSO and ACO algorithms.</title>
  			</caption>
   <table>
      <thead>
		 <tr>
		 	<th>Method</th>
			<th>Selected Features</th>
			<th>Number of Selected Features</th>			
		</tr>		 
      </thead>
      <tbody>
	     <tr>
			<td>BPSO</td>	
			<td>134, 162, 186, 320, 321, 333, 342, 351, 352, 353</td>
			<td>10</td>			
         </tr>        
         <tr>
            <td>ACO</td>	
			<td>352, 381, 419, 353, 342</td>
			<td>5</td>													
         </tr>
	</tbody>
	</table>
	</table-wrap>
	<table-wrap position="float" id="t3">
	<label>Table 3.</label>
  			<caption>
  				<title>PROSITE patterns selected in more than 80% of the runs performed by the ACO algorithm.</title>
  			</caption>
   <table>
      <thead>
		 <tr>
		 	<th>Feature / PROSITE pattern ID</th>
			<th>Name</th>				
		</tr>		 
      </thead>
      <tbody>
	     <tr>
			<td><italic>F</italic><sub>342</sub>/ps00410</td>	
			<td>Dynamin</td>					
         </tr>        
         <tr>
            <td><italic>F</italic><sub>352</sub>/ps00236</td>				
			<td>Neurotransmitter-gated ion-channel</td>													
         </tr>
		 <tr>
            <td><italic>F</italic><sub>353</sub>/ps00237</td>				
			<td>Rhodopsin-like GPCR superfamily</td>													
         </tr>
	</tbody>
	</table>
	</table-wrap>
	<table-wrap position="float" id="t4">
	<label>Table 4.</label>
  			<caption>
  				<title>Comparison of obtained results for ACO and standard binary PSO using nearest neighbor.</title>
  			</caption>
   <table>
      <thead>
		 <tr>
		 	<th>&nbsp;</th>
			<th colspan="3">Using all the 443 original feature</th>
			<th colspan="4">Standard Binary PSO algorithm</th>
			<th colspan="4">Proposed ACO Algorithm</th>			
		</tr>
		<tr>
		   <th>Fold</th>
		   <th>TPR</th>
		   <th>TNR</th>
		   <th>TPR &times; TNR</th>
		   <th>TPR</th>
		   <th>TNR</th>
		   <th>TPR &times; TNR</th>
		   <th>No. of selected features</th>
		   <th>TPR</th>
		   <th>TNR</th>
		   <th>TPR &times; TNR</th>
		   <th>No. of selected features</th>		   
		</tr>		 
      </thead>
      <tbody>
	     <tr>
			<td>1</td>	
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>									
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>16</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>6</td>
         </tr>        
         <tr>
            <td>2</td>	
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>									
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>19</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>2</td>										
         </tr>
		 <tr>
            <td>3</td>	
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>									
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>21</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>4</td>										
         </tr>
		 <tr>
            <td>4</td>	
			<td>0.73</td>
			<td>1.00</td>
			<td>0.73</td>									
			<td>0.76</td>
			<td>1.00</td>
			<td>0.76</td>
			<td>14</td>
			<td>0.73</td>
			<td>1.00</td>
			<td>0.73</td>
			<td>3</td>										
         </tr>
		 <tr>
            <td>5</td>	
			<td>0.00</td>
			<td>1.00</td>
			<td>0.00</td>									
			<td>0.69</td>
			<td>0.92</td>
			<td>0.63</td>
			<td>17</td>
			<td>0.73</td>
			<td>0.96</td>
			<td>0.70</td>
			<td>5</td>										
         </tr>
		 <tr>
            <td>6</td>	
			<td>0.00</td>
			<td>1.00</td>
			<td>0.00</td>									
			<td>0.65</td>
			<td>0.96</td>
			<td>0.62</td>
			<td>18</td>
			<td>0.88</td>
			<td>0.99</td>
			<td>0.87</td>
			<td>5</td>										
         </tr>
		 <tr>
            <td>7</td>	
			<td>0.92</td>
			<td>1.00</td>
			<td>0.92</td>									
			<td>0.88</td>
			<td>1.00</td>
			<td>0.88</td>
			<td>11</td>
			<td>0.92</td>
			<td>1.00</td>
			<td>0.92</td>
			<td>9</td>										
         </tr>
		 <tr>
            <td>8</td>	
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>									
			<td>0.69</td>
			<td>1.00</td>
			<td>0.69</td>
			<td>15</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>6</td>										
         </tr>
		 <tr>
            <td>9</td>	
			<td>0.73</td>
			<td>1.00</td>
			<td>0.73</td>									
			<td>0.73</td>
			<td>1.00</td>
			<td>0.73</td>
			<td>9</td>
			<td>0.73</td>
			<td>1.00</td>
			<td>0.73</td>
			<td>3</td>										
         </tr>
		 <tr>
            <td>10</td>	
			<td>0.42</td>
			<td>1.00</td>
			<td>0.42</td>									
			<td>0.42</td>
			<td>1.00</td>
			<td>0.42</td>
			<td>14</td>
			<td>0.42</td>
			<td>1.00</td>
			<td>0.42</td>
			<td>8</td>										
         </tr>
		 <tr>
            <td><bold>AVG</bold></td>	
			<td><bold>0.68</bold></td>
			<td><bold>1.00</bold></td>
			<td><bold>0.68</bold></td>									
			<td><bold>0.78</bold></td>
			<td><bold>0.99</bold></td>
			<td><bold>0.77</bold></td>
			<td><bold>15.4</bold></td>
			<td><bold>0.84</bold></td>
			<td><bold>0.99</bold></td>
			<td><bold>0.83</bold></td>
			<td><bold>5.1</bold></td>										
         </tr>
	</tbody>
	</table>
	</table-wrap>
	<table-wrap position="float" id="t5">
	<label>Table 5.</label>
  			<caption>
  				<title>Comparison of obtained results for ACO and standard binary PSO using na&iuml;ve Bayes classifier.</title>
  			</caption>
   <table>
      <thead>
		 <tr>
		 	<th>&nbsp;</th>
			<th colspan="3">Using all the 443 original feature</th>
			<th colspan="4">Standard Binary PSO algorithm</th>
			<th colspan="4">Proposed ACO Algorithm</th>			
		</tr>
		<tr>
		   <th>Fold</th>
		   <th>TPR</th>
		   <th>TNR</th>
		   <th>TPR &times; TNR</th>
		   <th>TPR</th>
		   <th>TNR</th>
		   <th>TPR &times; TNR</th>
		   <th>No. of selected features</th>
		   <th>TPR</th>
		   <th>TNR</th>
		   <th>TPR &times; TNR</th>
		   <th>No. of selected features</th>		   
		</tr>		 
      </thead>
      <tbody>
	     <tr>
			<td>1</td>	
			<td>0.92</td>
			<td>1.00</td>
			<td>0.92</td>									
			<td>0.88</td>
			<td>1.00</td>
			<td>0.88</td>
			<td>12</td>
			<td>0.92</td>
			<td>1.00</td>
			<td>0.92</td>
			<td>6</td>
         </tr>        
         <tr>
            <td>2</td>	
			<td>0.69</td>
			<td>1.00</td>
			<td>0.69</td>									
			<td>0.69</td>
			<td>1.00</td>
			<td>0.69</td>
			<td>13</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>4</td>										
         </tr>
		 <tr>
            <td>3</td>	
			<td>0.73</td>
			<td>1.00</td>
			<td>0.73</td>									
			<td>0.73</td>
			<td>1.00</td>
			<td>0.73</td>
			<td>11</td>
			<td>0.92</td>
			<td>1.00</td>
			<td>0.92</td>
			<td>4</td>										
         </tr>
		 <tr>
            <td>4</td>	
			<td>0.00</td>
			<td>1.00</td>
			<td>0.00</td>									
			<td>0.88</td>
			<td>0.96</td>
			<td>0.84</td>
			<td>9</td>
			<td>0.88</td>
			<td>0.96</td>
			<td>0.84</td>
			<td>5</td>										
         </tr>
		 <tr>
            <td>5</td>	
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>									
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>11</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>4</td>										
         </tr>
		 <tr>
            <td>6</td>	
			<td>0.69</td>
			<td>0.96</td>
			<td>0.66</td>									
			<td>0.69</td>
			<td>0.96</td>
			<td>0.66</td>
			<td>13</td>
			<td>0.73</td>
			<td>0.96</td>
			<td>0.70</td>
			<td>5</td>										
         </tr>
		 <tr>
            <td>7</td>	
			<td>0.42</td>
			<td>1.00</td>
			<td>0.42</td>									
			<td>0.42</td>
			<td>1.00</td>
			<td>0.42</td>
			<td>8</td>
			<td>0.42</td>
			<td>1.00</td>
			<td>0.42</td>
			<td>5</td>										
         </tr>
		 <tr>
            <td>8</td>	
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>									
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>7</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>2</td>										
         </tr>
		 <tr>
            <td>9</td>	
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>									
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>15</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>1.00</td>
			<td>3</td>										
         </tr>
		 <tr>
            <td>10</td>	
			<td>0.76</td>
			<td>1.00</td>
			<td>0.76</td>									
			<td>0.76</td>
			<td>1.00</td>
			<td>0.76</td>
			<td>10</td>
			<td>0.76</td>
			<td>1.00</td>
			<td>0.76</td>
			<td>4</td>										
         </tr>
		 <tr>
            <td><bold>AVG</bold></td>	
			<td><bold>0.72</bold></td>
			<td><bold>0.99</bold></td>
			<td><bold>0.71</bold></td>									
			<td><bold>0.80</bold></td>
			<td><bold>0.99</bold></td>
			<td><bold>0.79</bold></td>
			<td><bold>10.9</bold></td>
			<td><bold>0.86</bold></td>
			<td><bold>0.99</bold></td>
			<td><bold>0.85</bold></td>
			<td><bold>4.2</bold></td>										
         </tr>
	 </tbody>
	 </table>
	 </table-wrap>
	 </floats-wrap>		  					  						 
</article>