<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
	<front>
		<journal-meta>
			<journal-id journal-id-type="nlm-ta">J Comput Sci Syst Biol</journal-id>
			<journal-id journal-id-type="publisher-id">opg</journal-id>
			<journal-title-group>						
			<journal-title>Journal of Computer Science &amp; Systems Biology</journal-title>
			</journal-title-group>			 
			<issn pub-type="epub">0974-7230</issn>
			<publisher>
				<publisher-name>OMICS Publishing Group</publisher-name>
				<publisher-loc>India, USA</publisher-loc>
			</publisher>
		</journal-meta>
		<article-meta>	
			<article-id pub-id-type="doi">10.4172/jcsb.1000045</article-id>		
			<article-id pub-id-type="publisher-id">000063</article-id>
			<article-categories>
				<subj-group subj-group-type="heading">
					<subject>Research Article</subject>
				</subj-group>
				<subj-group subj-group-type="Discipline">
					<subject>Biochemistry</subject>
				</subj-group>
				<subj-group subj-group-type="System Taxonomy">
					<subject>Proteomics</subject>
					<subject>Bioinformatics</subject>
					<subject>Genomics</subject>
					<subject>Transcriptomics</subject>
					<subject>Biomarkers</subject>
				</subj-group>
			</article-categories>
			<title-group>
				<article-title>Prokaryotic and Eukaryotic Non-membrane Proteins have Biased Amino Acid Distribution</article-title>
			</title-group>
			<contrib-group>
				<contrib contrib-type="author">
					<name>
						<surname>Gaur</surname>
						<given-names>Rajneesh Kumar</given-names>						
					</name>																			
				</contrib>													
			</contrib-group>
			<aff>Bioinformatics Infrastructure Facility, Jamia Hamdard (Hamdard University), Hamdard Nagar, New Delhi, India - 110062</aff>			
			<author-notes>
				<corresp id="cor1">&ast; To whom correspondence should be addressed: Rajneesh Kumar Gaur, Bioinformatics Infrastructure Facility, Jamia Hamdard (Hamdard University), Hamdard Nagar, New Delhi, India - 110062, Tel: +91 9990290384; E-mail: <email>meetgaur@gmail.com</email></corresp>
			</author-notes>
			<pub-date pub-type="collection">
			     <month>12</month>
				 <year>2009</year>
			</pub-date>
			<pub-date pub-type="epub">
				<day>27</day>
				<month>12</month>
				<year>2009</year>
			</pub-date>			
			<volume>2</volume>
			<issue>6</issue>
			<fpage>298</fpage>
			<lpage>299</lpage>
			<history>
			<date date-type="received">
			     <day>30</day>
				 <month>09</month>
				 <year>2009</year>
			</date>
			<date date-type="accepted">
			      <day>27</day>
				  <month>12</month>
				  <year>2009</year>
			</date>
			</history>
			<permissions>			 
			<copyright-statement><bold>Copyright:</bold> &copy; 2009 Gaur RK.</copyright-statement>
			<copyright-year>2009</copyright-year>
			<license license-type="open access">
			 <license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.</license-p>
			 </license>
			 </permissions>			
			<abstract>
				<p>Proteins constitute the important constituent of the cellular machinery. The comparative analysis of non-membrane proteins (nMPs) between prokaryotes and eukaryotes carried out to determine the biasedness in amino acid distribution. On comparison, the results revealed that 'Ala' is the dominant amino acid in prokaryotic nMPs while 'Lys, Ser and Cys' are the dominant amino acids in eukaryotic nMPs.</p>
			</abstract>
			<kwd-group>
				<kwd>Non-membrane proteins</kwd>
				<kwd>Amino acid composition</kwd>
				<kwd>Prokaryotes</kwd>
				<kwd>Eukaryotes</kwd>																	
			</kwd-group>			
		</article-meta>
	</front>	
	<body>
		<sec>
			<title>Introduction</title>
				<p>Proteins constitute about 50% of the dry weight of most cells and are the most structurally complex macromolecules known. Proteins can be classified in different manner but for the purpose of this study we classified them as membrane (part of either cellular or organelle membrane; MPs) and non-membrane (located outside the membrane; nMPs) proteins. Amino acids are the building block of a protein and their composition determines the overall properties and stability of a protein. Many previous studies have shown how amino acid composition can be successfully applied to protein sequence analysis, including prediction of structural class (<xref ref-type="bibr" rid="r16">Zhang et al., 1992</xref>), discrimination of intra- and extra cellular proteins (<xref ref-type="bibr" rid="r5">Nakashima et al., 1994</xref>), prediction of sub-cellular location (<xref ref-type="bibr" rid="r1">Cedano et al., 1997</xref>). It was suggested that composition differences are a consequence of different requirements for protein folding, stability and transportation.</p>
				<p>The recent increase in the number of whole genome sequences has made the analysis of the corresponding proteomes possible. So far the amino acid composition of both the prokaryotic and eukaryotic proteomic databases have been explored separately for different purposes such as determination of sequence length (<xref ref-type="bibr" rid="r4">Gerstein, 1998a</xref>), identification of conserved sequences (<xref ref-type="bibr" rid="r9">Sobolevsky et al., 2005</xref>); elucidation of simple sequences (Subramanyam et al., 2006) etc. However, till now the comparative analysis of their non-membrane proteins (nMPs) have not been carried out to determine the overall amino acid compositional differences. This computational study is performed to develop the amino acid distribution of proteins as a tool to identify the proteins frequently undergo mutations and largely responsible for the pathogenicity of the organism.</p>				
				<fig id="g1">
					<label>Figure 1</label>
					<caption>
						<title>Histogram showing the overall amino acid composition of prokaryotic (black bars) and eukaryotic (white bars) nMPs. The amino acids are arranged in decreasing order of hydrophobicity. Pro: Prokaryotic nMPs; Euk: Eukaryotic nMPs.</title>																		
					</caption>
					<graphic xlink:href="JCSB-02-298-g001.tif"/>
				</fig>											
		</sec>					
				<sec sec-type="material||methods">
						<title>Methodology</title>											      						  
						<p>The dataset was curated manually from the sequences extracted from PSORT (<xref ref-type="bibr" rid="r8">Rey et al., 2005</xref>), eSLDB (<xref ref-type="bibr" rid="r6">Pierleoni et al., 2007</xref>) and RefSeq (<xref ref-type="bibr" rid="r7">Pruittet et al., 2005</xref>) databases. Only the experimentally annotated entries were extracted from PSORT database. From the RefSeq database, we used microbial (microbial1.protein.faa.gz; 05/11/2009) and eukaryotic (vertebrate_mammalian1.protein. faa.gz; 05/11/2009 &amp; vertebrate_other1.protein.faa.gz; 05/10/2009) sequence release files for construction of the experimental dataset. Protein sequences flagged as putat ive, hypothet ical , potent ial , uncharacterized, similar to the predicted protein, membrane, porin, receptor are deleted from the initially downloaded RefSeq sequence release files in the preparation of experimental dataset. The prokaryotic sequence dataset was created by merging the sequence entries from PSORT db and refseq dataset after appropriate deletions. Similarly, the eukaryotic dataset was prepared after deleting and merging the sequence entries from eSLDB and refseq dataset.</p>
						<p>The entire dataset used for computing the composition of 20 amino acid residues comprised of prokaryotic (63644) and eukaryotic (88400) nMP sequences. The amino acid composition for the prepared datasets was computed using the number of amino acids of each type and the total number of residues. It is defined as Residue composition (%) (r) = &Sigma;n<sub>r</sub>/N X100 (1) where 'r' stands for any one of the 20 amino acid residue. &Sigma;n<sub>r</sub> is the total number of residue of each type and N is the total number of residues in the dataset.</p>							
				</sec>
				<sec>
					<title>Results and Discussion</title>					   
					<p>The amino acid compositional distribution between prokaryotic and eukaryotic nMPs was computed using eq. (1). The prokaryotic nMPs shows the dominant occurrence of a non-polar amino acid 'Ala' (&sigma; = 0.45) while the eukaryotic nMPs predominantly possess the polar amino acids 'Lys' (&sigma; = 0.66), 'Ser' (&sigma; = 0.60) and 'Cys' (&sigma; = 0.29) (<xref ref-type="fig" rid="g1">Figure 1</xref>). In prokaryotic nMPs, the high frequency of short side-chained non-polar aliphatic amino acid 'Ala' may be due to various possibilities such as its over-representation in highly expressed proteins (<xref ref-type="bibr" rid="r13">Tats et al., 2006</xref>), its role in determining the cleavage of N-terminal formyl methionine (<xref ref-type="bibr" rid="r10">Solbiati et al., 1999</xref>), its role in assisting the entrance of the nascent peptide chain into the ribosomal tunnel (<xref ref-type="bibr" rid="r14">Tenson et al., 2002</xref>) and in helix-helix packing (<xref ref-type="bibr" rid="r3">Eyre et al., 2004</xref>). Though 'Ala' might perform the similar functions in both prokaryotic and eukaryotic nMPs but its higher frequency in nMPs probably related to the higher proportion of prokaryotic helical nMPs.</p>
					<p>The eukaryotes show the high occurrence of positively charged polar residue 'Lys' in their nMPs repertoire. This positively charged residue helps in the secretion of proteins through the membrane via interaction with export machinery and signal recognition particles (<xref ref-type="bibr" rid="r15">vonHeijne, 1984</xref>). The overabundance of 'Ser' in eukaryotic nMPs may be due to their ability to form H-bonds and stabilizing the helices (Subramaniam et al., 2006). In particular, the two-fold higher 'Cys' of eukaryotic nMPs compared to prokaryotic nMPs most probably compensates for their lower hydrophobicity (<xref ref-type="bibr" rid="r2">D'Onofrio et al., 1999</xref>).</p>
					</sec>																																	
				</body>
				<back>
				<ack> 
			<p>I express my gratitude to the Council of Scientific and Industrial Research (CSIR), New Delhi, India for granting me the Senior Research Associateship. I am also thankful to Dr. Sayeed Ahmed, Faculty of Pharmacy, Jamia Hamdard University, New Delhi, India for extending his computational facility.</p>			
		</ack>						
		<ref-list>
			<title>References</title>
				<ref id="r1">
				 <label>1</label>
                 <element-citation publication-type="journal">							 
							<name>
								<surname>Cedano</surname>
								<given-names>J</given-names>
							</name>
							<name>
								<surname>Aloy</surname>
								<given-names>P</given-names>								
							</name>		
							<name>
								<surname>Perez-Pons</surname>
								<given-names>JA</given-names>								
							</name>
							<name>
								<surname>Querol</surname>
								<given-names>E</given-names>								
							</name>																																		
							<article-title>Relation between amino acid composition and cellular location of proteins</article-title>
							<source>J Mol Biol</source>	
							<year>1997</year>
							<volume>266</volume>
							<fpage>594</fpage>
							<lpage>600</lpage>						
				</element-citation>
				</ref>
				<ref id="r2">
				<label>2</label>
				<element-citation publication-type="journal">							 
							<name>
								<surname>D'Onofrio</surname>
								<given-names>G</given-names>
							</name>
							<name>
								<surname>Jabbari</surname>
								<given-names>K</given-names>								
							</name>		
							<name>
								<surname>Musto</surname>
								<given-names>H</given-names>								
							</name>
							<name>
								<surname>Bernardi</surname>
								<given-names>G</given-names>								
							</name>																																		
							<article-title>The correlation of protein hydropathy with the base composition of coding sequences</article-title>
							<source>Gene</source>	
							<year>1999</year>
							<volume>238</volume>
							<fpage>3</fpage>
							<lpage>14</lpage>							
			 	</element-citation>
				</ref>
				<ref id="r3">
				<label>3</label>
				<element-citation publication-type="journal">							 
							<name>
								<surname>Eyre</surname>
								<given-names>TA</given-names>
							</name>
							<name>
								<surname>Partridge</surname>
								<given-names>L</given-names>								
							</name>		
							<name>
								<surname>Thornton</surname>
								<given-names>JM</given-names>								
							</name>																																
							<article-title>Computational analysis of {alpha}-helical membrane protein structure: implications for the prediction of 3D structural models</article-title>
							<source>Protein Eng Des Sel</source>	
							<year>2004</year>
							<volume>17</volume>
							<fpage>613</fpage>
							<lpage>624</lpage>							
			 	</element-citation>
				</ref>
				<ref id="r4">
				<label>4</label>
				<element-citation publication-type="journal">							 
							<name>
								<surname>Gerstein</surname>
								<given-names>M</given-names>
							</name>																																									
							<article-title>How representative are the known structures of the proteins in a complete genome? A comprehensive structural census</article-title>
							<source>Fold Des</source>	
							<year>1998a</year>
							<volume>3</volume>
							<fpage>497</fpage>
							<lpage>512</lpage>														
			 	</element-citation>
				</ref>
				<ref id="r5">
				<label>5</label>
				<element-citation publication-type="journal">							 
							<name>
								<surname>Nakashima</surname>
								<given-names>H</given-names>
							</name>
							<name>
								<surname>Nishikawa</surname>
								<given-names>K</given-names>								
							</name>																																		
							<article-title>Discrimination of intracellular and extracellular proteins using amino acid composition and residue-pair frequencies</article-title>
							<source>J Mol Biol</source>	
							<year>1994</year>
							<volume>238</volume>
							<fpage>54</fpage>
							<lpage>61</lpage>
			 	</element-citation>
				</ref>
				<ref id="r6">
				<label>6</label>
				<element-citation publication-type="journal">							
							<name>
								<surname>Pierleoni</surname>
								<given-names>A</given-names>								
							</name>
							<name>
								<surname>Martelli</surname>
								<given-names>PL</given-names>								
							</name>
							<name>
								<surname>Fariselli</surname>
								<given-names>P</given-names>								
							</name>
							<name>
								<surname>Casadio</surname>
								<given-names>R</given-names>								
							</name>																																										
							<article-title>eSLDB: Eukaryotic subcellular localization databse</article-title>
							<source>Nucleic Acids Res</source>
							<year>2007</year>
							<volume>35</volume>
							<fpage>D208</fpage>
							<lpage>212</lpage>														
				</element-citation>
				</ref>
				<ref id="r7">
				<label>7</label>
				<element-citation publication-type="journal">							
							<name>
								<surname>Pruitt</surname>
								<given-names>KD</given-names>								
							</name>
							<name>
								<surname>Tatusova</surname>
								<given-names>T</given-names>								
							</name>
							<name>
								<surname>Maglott</surname>
								<given-names>DR</given-names>								
							</name>																																																	
							<article-title>NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins</article-title>
							<source>Nucleic Acids Res</source>
							<year>2005</year>
							<volume>33</volume>
							<fpage>D501</fpage>
							<lpage>504</lpage>							 
			 	</element-citation>
				</ref>
				<ref id="r8">
				<label>8</label>
				<element-citation publication-type="journal">							
							<name>
								<surname>Rey</surname>
								<given-names>S</given-names>								
							</name>
							<name>
								<surname>Acab</surname>
								<given-names>M</given-names>								
							</name>
							<name>
								<surname>Gardy</surname>
								<given-names>JL</given-names>								
							</name>
							<name>
								<surname>Laird</surname>
								<given-names>MR</given-names>								
							</name>
							<name>
								<surname>deFays</surname>
								<given-names>K</given-names>								
							</name><etal/>																																
 							<article-title>PSORTdb: A Database of Subcellular Localizations for Bacteria</article-title>
							<source>Nucleic Acids Res</source>
							<year>2005</year>
							<volume>33</volume>
							<fpage>D164</fpage>
							<lpage>168</lpage>		
			 	</element-citation>
				</ref>
				<ref id="r9">
				<label>9</label>
				<element-citation publication-type="journal">							
							<name>
								<surname>Sobolevsky</surname>
								<given-names>Y</given-names>								
							</name>
							<name>
								<surname>Trifonov</surname>
								<given-names>EN</given-names>								
							</name>																																
 							<article-title>Conserved sequences of prokaryotic proteomes and their computational age</article-title>
							<source>J Mol Evol</source>
							<year>2005</year>
							<volume>61</volume>
							<fpage>591</fpage>
							<lpage>596</lpage> 														
			 	</element-citation>
				</ref>	
				<ref id="r10">
				<label>10</label>
				<element-citation publication-type="journal">							
							<name>
								<surname>Solbiati</surname>
								<given-names>J</given-names>								
							</name>
							<name>
								<surname>Chapman-Smith</surname>
								<given-names>A</given-names>								
							</name>
							<name>
								<surname>Miller</surname>
								<given-names>JL</given-names>								
							</name>
							<name>
								<surname>Miller</surname>
								<given-names>CG</given-names>								
							</name>
							<name>
								<surname>Cronan</surname>
								<given-names>JEJ</given-names>								
							</name>																																
 							<article-title>Processing of the N termini of nascent polypeptide chains requires deformylation prior to methionine removal</article-title>
							<source>J Mol Biol</source>
							<year>1999</year>
							<volume>290</volume>
							<fpage>607</fpage>
							<lpage>614</lpage>		
			 	</element-citation>
				</ref>	
				<ref id="r11">
				<label>11</label>
				<element-citation publication-type="journal">							
							<name>
								<surname>Subramaniam</surname>
								<given-names>S</given-names>								
							</name>
							<name>
								<surname>Henderson</surname>
								<given-names>R</given-names>								
							</name>																																							
 							<article-title>Molecular mechanism of vectorial proton translocation by bacteriorhodopsin</article-title>
							<source>Nature</source>
							<year>2000</year>
							<volume>406</volume>
							<fpage>653</fpage>
							<lpage>657</lpage>		
			 	</element-citation>
				</ref>	
				<ref id="r12">
				<label>12</label>
				<element-citation publication-type="journal">							
							<name>
								<surname>Subramanyam</surname>
								<given-names>MB</given-names>								
							</name>
							<name>
								<surname>Gnanamani</surname>
								<given-names>M</given-names>								
							</name>
							<name>
								<surname>Ramachandran</surname>
								<given-names>S</given-names>								
							</name>																																							
 							<article-title>Simple sequence proteins in prokaryotic proteome</article-title>
							<source>BMC Genomics</source>
							<year>2006</year>
							<volume>7</volume>
							<fpage>141</fpage>																
			 	</element-citation>
				</ref>	
				<ref id="r13">
				<label>13</label>
				<element-citation publication-type="journal">							
							<name>
								<surname>Tats</surname>
								<given-names>A</given-names>								
							</name>
							<name>
								<surname>Remm</surname>
								<given-names>M</given-names>								
							</name>
							<name>
								<surname>Tenson</surname>
								<given-names>T</given-names>								
							</name>																																							
 							<article-title>Highly expressed proteins have an increased frequency of alanine in the second amino acid position</article-title>
							<source>BMC Genomics</source>
							<year>2006</year>
							<volume>7</volume>
							<fpage>28</fpage>																
			 	</element-citation>
				</ref>
				<ref id="r14">
				<label>14</label>
				<element-citation publication-type="journal">							
							<name>
								<surname>Tenson</surname>
								<given-names>T</given-names>								
							</name>
							<name>
								<surname>Ehrenberg</surname>
								<given-names>M</given-names>								
							</name>																																							
 							<article-title>Regulatory nascent peptides in the ribosomal tunnel</article-title>
							<source>Cell</source>
							<year>2002</year>
							<volume>108</volume>
							<fpage>591</fpage>
							<lpage>594</lpage>		
			 	</element-citation>
				</ref>
				<ref id="r15">
				<label>15</label>
				<element-citation publication-type="journal">							
							<name>
								<surname>vonHeijne</surname>
								<given-names>G</given-names>								
							</name>																																														
 							<article-title>Analysis of the distribution of charged residues in the N-terminal region of signal sequences: implications for protein export in prokaryotic and eukaryotic cells</article-title>
							<source>EMBO J</source>
							<year>1984</year>
							<volume>3</volume>
							<fpage>2315</fpage>
							<lpage>2318</lpage>		
			 	</element-citation>
				</ref>
				<ref id="r16">
				<label>16</label>
				<element-citation publication-type="journal">							
							<name>
								<surname>Zhang</surname>
								<given-names>CT</given-names>								
							</name>
							<name>
								<surname>Chou</surname>
								<given-names>KC</given-names>								
							</name>																																														
 							<article-title>An optimization approach to predicting protein structural class from amino acid composition</article-title>
							<source>Protein Sci</source>
							<year>1992</year>
							<volume>1</volume>
							<fpage>401</fpage>
							<lpage>408</lpage>		
			 	</element-citation>
				</ref>										
        </ref-list>
		<glossary>
		  <title>Abbreviations</title>
		    <def-list>
			  <def-item>
			    <term>MPs</term>
				<def>
				  <p>Membrane Proteins</p>
				</def>
			  </def-item>
			  <def-item>
			    <term>nMPs</term>
				<def>
				  <p>non-membrane proteins</p>
				</def>
			  </def-item>
			</def-list>   	  
		</glossary>										
		</back>				
</article>