<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
	<front>
		<journal-meta>
			<journal-id journal-id-type="nlm-ta">J Canc Sci Ther</journal-id>
			<journal-id journal-id-type="publisher-id">opg</journal-id>
			<journal-title-group>						
			<journal-title>Journal of Cancer Science &amp; Therapy</journal-title>
			</journal-title-group>			 
			<issn pub-type="epub">1948-5956</issn>
			<publisher>
				<publisher-name>OMICS Publishing Group</publisher-name>
				<publisher-loc>India, USA</publisher-loc>
			</publisher>
		</journal-meta>
		<article-meta>	
			<article-id pub-id-type="doi">10.4172/1948-5956.1000004</article-id>		
			<article-id pub-id-type="publisher-id">000063</article-id>
			<article-categories>
				<subj-group subj-group-type="heading">
					<subject>Research Article</subject>
				</subj-group>
				<subj-group subj-group-type="Discipline">
					<subject>Biochemistry</subject>
				</subj-group>
				<subj-group subj-group-type="System Taxonomy">
					<subject>Proteomics</subject>
					<subject>Bioinformatics</subject>
					<subject>Genomics</subject>
					<subject>Transcriptomics</subject>
					<subject>Biomarkers</subject>
				</subj-group>
			</article-categories>
			<title-group>
				<article-title>Model-Based Background Correction (MBCB): R Methods and GUI for Illumina Bead-array Data</article-title>
			</title-group>
			<contrib-group>
				<contrib contrib-type="author">
					<name>
						<surname>D. Allen</surname>
						<given-names>Jeffrey</given-names>
					</name>	
					<xref ref-type="aff" rid="a1">1</xref>
					<xref ref-type="aff" rid="a2">2</xref>								
				</contrib>
				<contrib contrib-type="author">
					<name>
						<surname>Chen</surname>
						<given-names>Min</given-names>
					</name>		
					<xref ref-type="aff" rid="a3">3</xref>										
				</contrib>	
				<contrib contrib-type="author">
					<name>
						<surname>Xie</surname>
						<given-names>Yang</given-names>
					</name>	
					<xref ref-type="aff" rid="a2">2</xref>
					<xref ref-type="aff" rid="a4">4</xref>									
					<xref ref-type="corresp" rid="cor1">&ast;</xref>									
				</contrib>	
			</contrib-group>
			<aff id="a1"><label>1</label>Department of Computer Science and Engineering, Southern Methodist University, Dallas, TX, 75205, USA</aff>
			<aff id="a2"><label>2</label>Division of Biostatistics, Department of Clinical Sciences, University of Texas Southwestern Medical Center, 5323 Harry Hines Blvd, Dallas, TX, 75390, USA</aff>
			<aff id="a3"><label>3</label>Department of Epidemiology and Public Health, Yale University, 300 George Street, Suite 503, New Haven, CT 06510, USA</aff>
			<aff id="a4"><label>4</label>Harold C. Simmons Comprehensive Cancer Center University of Texas, Southwestern Medical Center, 5323 Harry Hines Blvd, Dallas, TX, 75390, USA</aff>			
			<author-notes>
				<corresp id="cor1">&ast; To whom correspondence should be addressed: Yang Xie, Division of Biostatistics, Department of Clinical Sciences, University of Texas Southwestern Medical Center, 5323 Harry Hines Blvd, Dallas, TX, 75390, USA, E-mail: <email>yang.xie@utsouthwestern.edu</email></corresp>
			</author-notes>
			<pub-date pub-type="collection">
			     <month>11</month>
				 <year>2009</year>
			</pub-date>
			<pub-date pub-type="epub">
				<day>25</day>
				<month>11</month>
				<year>2009</year>
			</pub-date>			
			<volume>1</volume>
			<issue>1</issue>
			<fpage>025</fpage>
			<lpage>027</lpage>
			<history>
			<date date-type="received">
			     <day>01</day>
				 <month>11</month>
				 <year>2009</year>
			</date>
			<date date-type="accepted">
			      <day>25</day>
				  <month>11</month>
				  <year>2009</year>
			</date>
			</history>
			<permissions>			 
			<copyright-statement><bold>Copyright:</bold> &copy; 2009 Allen JD, et al.</copyright-statement>
			<copyright-year>2009</copyright-year>
			<license license-type="open access">
			 <license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.</license-p>
			 </license>
			 </permissions>			
			<abstract>
				<p><bold>Summary:</bold> Illumina BeadArray platform (Illumina Inc.) is playing an increasing role in cancer research. MBCB, an R package designed for use on Illumina Bead-Array data, allows for microarray data to be pre-processed through various model-based statistical methods. These model-based background-correction methods have proven to be a significant improvement over the traditional methods provided by Illumina in their BeadStudio software. MBCB accepts the summarized bead-type data; the data can then be normalized and background-corrected in a statistically-efficient manner. When compared to the popular Robust multi-array (RMA) background correction approach and the default, Illuminaprovided background-correction method, MBCB has shown to lead to more precise determination of gene expression and better biological interpretation of Illumina BeadArray data. The software developed will facilitate molecular biomedical - especially cancer - research.</p>
				<p><bold>Availability:</bold> This package will soon be available from Bioconductor. Instructions for use are included with the package.</p>
			</abstract>
			<kwd-group>
				<kwd>Background correction</kwd>				
				<kwd>Microarray</kwd>				
				<kwd>BeadArray</kwd>									
			</kwd-group>			
		</article-meta>
	</front>	
	<body>
	       <sec>
		      <title>Introduction</title>		
				<p>Illumina have produced a novel microarray platform - BeadArray - for use in multiple environments: Gene expression studies, CGH, and SNP, among others. Illumina expression BeadArrays often generate high quality data with relatively low cost and less RNA sample input. These features make BeadArray an increasingly popular microarray platform.</p>				
						<p>One distinguishing feature of the BeadArray platform is that each array contains thousands of non-specific negative control bead types. These negative control beads offer great potential for controlling background noise.</p>
						<p>The background-correction method in BeadStudio software provided by Illumina Inc. did not use negative control beads efficiently. It takes the mean of all negative control beads and subtracts that value from all of the other beads. Unfortunately, this method tends to result in a large number of negative expression values which are typically discarded. In certain cases, more than half of the beads on the chip have been negative when using this method. Some studies (<xref ref-type="bibr" rid="r1">Barnes et al., 2005</xref>) have suggested that the pre-processing methods offered by Illumina will actually cause such massive data loss that the raw values should be used instead. However, using only the raw values has also shown to be problematic as significant data attenuation is observed when expression ratios are calculated between two expression-level data sets (<xref ref-type="bibr" rid="r2">Ding et al., 2008</xref>).</p>
						<p><xref ref-type="bibr" rid="r2">Ding et al., 2008</xref>) suggested an alternative model-based background-correction method to address this problem. <xref ref-type="bibr" rid="r3">Xie et al., (2009)</xref> proposed three different statistical methods to estimate the parameters in the model. We developed MBCB - an R package - to take advantage of these new methods. By using R, the package is inherently cross-platform, easily distributable and can be easily integrated into existing R tools.</p>
			</sec>
			<sec>
		      <title>Description</title>
			        <sec>
					<title>Input</title>			
						<p>The user provides MBCB with the summarized bead-type data. This summarization can be obtained through BeadStudio. Essentially, the file summarizes the raw, bead-level data and provides the average intensity and variance of each bead type.</p>
					</sec>
					<sec>
					<title>Background-correction</title>			
						<p>The primary contribution of this package is the ability to background-correct the given data in a statistically-efficient manner. The algorithms used no longer cause massive data attenuation; instead, they lead to more accurate measurement of gene expression levels.</p>
						<p>The user can select from a list of background-correction methods:</p>
						<p><bold>Maximum likelihood estimation:</bold> One of the more accurate methods built. Assuming Gaussian distribution for the noise term, MLE iteratively updates parameter estimates by making use of the non-specific beads on the microarray.</p>
						<p><bold>Gamma maximum likelihood estimation:</bold> This method is similar to Maximum Likelihood Estimation except that the noise term is assumed to have a Gamma distribution. It is preferred when the distribution of the non-specific negative control data is not symmetric.</p>
						<p><bold>Bayesian method:</bold> Is possibly more extensible than the others (because it allows for extra prior information), but has not consistently outperformed the other methods, despite being, by far, the most computationally-intensive.</p>
						<p><bold>Non-Parametric:</bold> Avoids the use of assumptions about the parameters of the model. This is the fastest and one of the most accurate methods.</p>
						<p><bold>Robust multi-array average:</bold> This method (modeled after the methods found in the Affymetrix package) can be used if the users could not provide negative control information.</p>
					</sec>
					<sec>
					<title>Normalization</title>
					<p>For convenience' sake, the package also provides the opportunity to normalize the data, if so desired, using either Quantile-Quantile normalization (from the <italic>affy</italic> package) or global (median) normalization. Obviously, normalization is not mandatory for this package. Users can also apply their own preferred normalization approach after using MBCB for background-correction.</p>
					</sec>
					<sec>
					<title>Output</title>
					<p>The files created by this package include the background-corrected data (one file per background-correction method used). In addition, the user is given a file detailing the parameter estimations for each correction method.</p>
					</sec>
					<sec>
					<title>Graphical user interface</title>
					<p>To ensure ease-of-use among non-technical audiences, we provide a graphical interface through which users can accomplish all that they could with the command-level functions. A screenshot is shown in <xref ref-type="fig" rid="g1">Figure 1</xref>. The user can browse to the data files, select one or multiple model-based background-correction methods from a list, then select normalization method(s). The user can then browse to the location at which they'd like to save the output.</p>
					</sec>
			</sec>
			<sec>
			  <title>Discussion</title>
			  <p>BeadArray technology has a great potential for molecular profiling especially for cancer research; however, because of the analysis issues surrounding the background-correction of the data, the platform has experienced significant data loss/attenuation. These new model-based methods have proven to be more accurate and efficient (<xref ref-type="bibr" rid="r2">Ding et al., 2008</xref>; <xref ref-type="bibr" rid="r3">Xie et al., 2009</xref>) and the user-friendly R package will simplify the use of this data processing approach.</p>			
						<fig id="g1">
					<label>Figure 1</label>
					<caption>
						<title>A screenshot of the background-correction options in MBCB.</title>																		
					</caption>
					<graphic xlink:href="JCST-01-025-g001.tif"/>
				</fig>
						<p>Because the package is written exclusively in R, it can be used on any Windows or UNIX-based operating system. Many previously-written tools surrounding this research have also been written in R. Due to the open-source nature of most such packages, the tools can easily be combined or modified to meet specific needs.</p>
			</sec>													
				</body>
				<back>	
				<ack> 
			<p>This work was support by NIH UL1 RR024982, NNJ05HD36G, and 1R21DA027592.</p>
		</ack>	
		<ref-list>
			<title>References</title>
				<ref id="r1">
				 <label>1</label>
                 <element-citation publication-type="journal">							 
							<name>
								<surname>Barnes</surname>
								<given-names>M</given-names>
							</name>		
							<name>
								<surname>Freudenberg</surname>
								<given-names>J</given-names>
							</name>	
							<name>
								<surname>Thompson</surname>
								<given-names>S</given-names>
							</name>
							<name>
								<surname>Aronow</surname>
								<given-names>B</given-names>
							</name>	
							<name>
								<surname>Pavlidis</surname>
								<given-names>P</given-names>
							</name>																									
							<article-title>Experimental comparison and cross-validation of the Affymetrix and Illumina gene expression analysis platforms</article-title>
							<source>Nucleic Acids Res</source>
							<year>2005</year>
							<volume>33</volume>
							<fpage>5914</fpage>
							<lpage>5923</lpage>
				</element-citation>
				</ref>
				<ref id="r2">
				<label>2</label>
				<element-citation publication-type="journal">							
							<name>
								<surname>Ding</surname>
								<given-names>LH</given-names>								
							</name>		
							<name>
								<surname>Xie</surname>
								<given-names>Y</given-names>								
							</name>
							<name>
								<surname>Park</surname>
								<given-names>S</given-names>								
							</name>		
							<name>
								<surname>Xiao</surname>
								<given-names>G</given-names>								
							</name>
							<name>
								<surname>Story</surname>
								<given-names>MD</given-names>								
							</name>																																																															 							<article-title>Enhanced identification and biological validation of differential gene expression via Illumina whole-genome expression arrays through the use of the model-based background correction methodology</article-title>
							<source>Nucleic Acids Res</source>
							<year>2008</year>
							<volume>36</volume>							
							<fpage>e58</fpage>							
			 	</element-citation>
				</ref>
				<ref id="r3">
				<label>3</label>
				<element-citation publication-type="journal">							 
							<name>
								<surname>Xie</surname>
								<given-names>Y</given-names>
							</name>		
							<name>
								<surname>Wang</surname>
								<given-names>X</given-names>
							</name>	
							<name>
								<surname>Story</surname>
								<given-names>M</given-names>
							</name>																								
							<article-title>Statistical methods of background cor rect ion for Illumina BeadAr ray data</article-title>
							<source>Bioinformatics</source>
							<year>2009</year>
							<volume>25</volume>
							<fpage>751</fpage>
							<lpage>757</lpage>							
			 	</element-citation>
				</ref>							
        </ref-list>									
		</back>				
</article>