From: keith steward [ksteward@bacterialbarcodes.com]
Sent: Wednesday, December 17, 2003 5:14 PM
To: Home
Subject: {Blocked Content} FW: rep-PCR data analysis
review and do background reading as needed to better comment
 
 -----Original Message-----
From: David.Renter@gov.ab.ca [mailto:David.Renter@gov.ab.ca]
Sent: Wednesday, December 17, 2003 3:14 PM
To: Alex Renwick; John.Berezowski@gov.ab.ca
Cc: Mimi Healy; jmanry@bacbarcodes.com; arenwick@bacbarcodes.com; ksteward@bacbarcodes.com
Subject: Re: rep-PCR data analysis


Alex,
Thank you for your detailed reply.  It sounds like you understand our different approach and have thought about some of these issues.  John and I have provided some clarifications and further questions below.  (- John please add to the discussion if I missed anything) I'm not sure we have reached a solution, but I believe we are on the right track.

Thanks again
Dave


David G. Renter, DVM, PhD
Veterinary Epidemiologist
Agri-Food Systems Branch
Food Safety Division
Alberta Agriculture, Food and Rural Development
Phone: 780/422-0275
Fax:  780/422-3438



"Alex  Renwick" <arenwick@bacbarcodes.com>

12/16/2003 11:35 AM

       
        To:        David.Renter@gov.ab.ca
        cc:        "Mimi Healy" <mhealy@bacbarcodes.com>, jmanry@bacbarcodes.com;arenwick@bacbarcodes.com;ksteward@bacbarcodes.com;;;;;
        Subject:        rep-PCR data analysis



Dr. Renter,

I am responding to your request to M. Healy for advice on analyzing rep-PCR
fingerprint data.  My own background is in statistics, and I have been
working with this data for two years.  This email outlines my view of the
problems and possible solutions.

You are correct in recognizing that your question is different from our usual
approach.  Our primary interest is in inferring genetic relationships, while
your question concerns relating genetic relatedness to non-genetic factors.


>> Yes exactly  

The difficulties in doing this with rep-PCR data are simply those common to
all analyses of data sets of high dimension.

 
The high dimension of the genetic data is essential our concern

Is modelling correlation coefficients a concern?

We have not worked out the
answers yet, but we have thought about the issues.

The aim is to explain the variation observed in rep-PCR fingerprints in terms
of various risk factors.  This calls for an analysis of variance (ANOVA)
approach.  ANOVA accounts for the total variation in a response variable (the
rep-PCR fingerprint, in this case) by quantifying the degree to which each
covariate explains the outcome.


>> This is essentially what we want to do, although we were planning on using a regression analysis rather than an ANOVA. We could use a linear, logistic, polynomial, etc. regression depending on what is appropriate (and why we asked about the 'data structure').  A colleague of ours has suggested a 'matrix regression' and he implied that the matrix could constitute the dependent variable.  We are not familar with this approach and have not been able to find references to this method (and are not sure it exists or would work for this application).
You point of Normality below is one reason we would avoid a 'regular' ANOVA.  In addition, we would like to describe the magnitude of an association rather than just test for significance (e.g. calculate Odds Ratios for risk factors that are significant in the final model).  Also keep in mind that due to the observational study design and sampling scheme we would have a very unbalanced number of observations for some risk factors.

Several issues keep us from a simple application ANOVA to rep-PCR fingerprint
data.  First, ANOVA is typically applied to univariate, or low dimensional,
data, while each of our "data points" is of dimension 850.  We can avoid this
problem by using the pair-wise distances as the (univariate) response
variable.  This, however, would introduce the complication that the set of
pair-wise distances violates the important assumption of mutual
independence.  Also, reliable interpretation of the simple application of
ANOVA requires that the response variable be Normally distributed (or, at
least in the neighborhood of Normality).

There are several possible approaches.  First, I refer you to a 1992
paper "Analysis of Molecular Variance Inferred from Metric Distances among
DNA Haplotypes:  Application to Human Mitochondrial DNA Restriction Data" by
Excoffier, et al., (Genetics 131: 479-491).  This paper describes a method of
assessing the significance of hierarchical population divisions based on pair-
wise distance data.  Significance is determined by a permutation test rather
than by the traditional F statistic in order to avoid relying on
inappropriate assumptions.  This method is implemented in the widely-used
genetic data analysis program "Arlequin" (http://lgb.unige.ch/arlequin/).  
There may be more recent implementations.  An internet search for "AMOVA"
yields a variety of resources.


The application of permutation tests is interesting and may be relevant.  However, I'm not sure how we would include (and deal with dependency in) multiple explantory variables (see variables below).(ie Model vs testing for associations)

Excoffier, et al., are concerned with spatial structure, so all explanatory
variables are binary, and different variables have a hierarchical, rather
than additive, structure.  Their method would require some modification to
apply to more general risk factors.

We would have hierarchical risk factors as well (e.g. some herds have multiple pens within herd and multiple animals within pens).  Most of are explanatory variables are categorical; except perhaps for temporal and spatial variables.  We are familiar with methods that deal with these types of explanatory variables in regression analyses.

Although there is an individual-based model underlying the Excoffier method,
the input data is pair-wise distances.  Another approach is to apply ANOVA to
individual-based statistics.  


This requires dimension reduction.  (This is
your suggestion of "a measure
(or series of measures) for each isolate that capture the data".)  The
challenge is to find statistics that reduce the information available in a
fingerprint to just one, or a few, numbers.


Yes exactly

One very simple way to reduce the dimension is to represent each fingerprint
by its distance from a "standard" profile.  (This, I think, is your idea of
a "rooted tree".)


Yes.
And if we used this option, we thought it may be useful to do a sensitivity analysis of the regression model (quite common in epidemiologic analyses of observation data).  In other words, develop a regression model for risk factors using a 'standard' or 'reference' strain to define the dependent variable relationship ('genetic relatedness').  Then change the strain used for comparisons to another 'standard(s)' which would change (slightly?) the dependent variable and may or may not impact which risk factors are significant.  This would tell us whether or not the 'standard' chosen for the genetic comparison effects the model (effects the significance/magnitude of explanatory variables).
 
 If all fingerprints are small variations of a single type
this might work well.  If a few types appear in the data, you might choose a
few standard types and represent each fingerprint by its distances to each
standard.  In this case, the data would no longer be univariate and a
multivariate analysis of variance would be appropriate.

A more common method of reducing dimension is to compute the principle
components decomposition of the data set, and to choose the first few
components for analysis.  This method optimally (in some sense) represents
the variation in the data in the fewest possible orthogonal components.

It seems like some method of dimension reduction for the dependent variable may be necessary.  I mentioned to Dr. Healy that we would require their input if this is done so that the 'outcome' variable still accurately capture the genetic relationships. Perhaps a principle components approach would be most appropriate for analysis, but I am not sure.
 
In any case, once the dimension is reduced to a manageable number, it will be
important to apply a transformation to the distances so that their

distribution is approximately Normal.  After this, the application of ANOVA
is direct.

Finally, let me refer you to a recent paper on a similar topic, Blackwood, et
al., (2003) "Terminal Restriction Fragment Length Polymorphism Data Analysis
for Quantitative Comparison of Microbial Communities", published in Applied
and Environmental Microbiology vol. 69, number 2.  The statistical analysis
is not deep, and it does not address your question directly, but it might
serve to generate some ideas about possible analyses.

In the above paragraphs I have addressed the problem of inferring the
importance of risk factors.  In your email you go on to say that you would
like to "group isolates based on the significant factors".  I'm not quite
sure what you mean by this.  


What I meant was: if our statistical model suggests that factorX is associated with genetic variability, then we would want to create dendrograms grouped by each level of factorX (for figures in a publication lets say).  In other words, only create dendrograms after the statistical analysis (i.e.we know which explanatory variables are associated with genetic variability)

More generally, we would like a better idea of
your final aim for the analysis.  It could be that our existing analysis in
the DiversiLab software may address your need.

Our final aim for the analysis is to apply 'standard' epidemiologic methods for statistical analysis of observational studies to these genetic data.  That means determining which 'model' best describes the variability in the genetic data given the explanatory variables we have collected.  This will tell us which factor(s) best fit the data and therefore the effect (significance and magnitude) of these factor on genetic variablity.

The main challenge as we see it is defining a dependent variable that captures the genetic relationships and lends itself to this type of statistical methods.

Alex Renwick

P. S.  As for the data itself, we are currently exporting it from the
Bionumerics software.  Is there a particular format (e.g., tab delimited)
that you prefer?


> ---------- Forwarded Message -----------
> From: David.Renter@gov.ab.ca
> To: tbittner@bacbarcodes.com, mhealy@bacbarcodes.com,
> kawillia@epi.umaryland.edu
> Sent: Fri, 12 Dec 2003 14:01:08 -0700
> Subject: Data & Analysis
>
> Hello all,
> John and I have tried to outline our intentions (FYI - by risk
> factors, we mean the species, farm, geographic location, time, etc.
> from where the isolates were recovered.):        In general, we
> would like to see which risk factors explain the genetic relatedness
> among isolates; rather than compare the genetic relatedness among
> isolates with different risk factors.  In other words, instead of
> generating hypotheses in advance about genetic relatedness that we
> test (e.g. isolates within farms are more/less similar than between
> farms) we would like to use the genetic data tell us which factors
> are important (e.g. does the 'farm' variable explain relatedness).
>  The difference between the two approaches may seem subtle, but
> statistically these are quite different.        The statistical
> model would be 'genetic relatedness' =(i.e. is explained by)  
> factor1 + factor2, etc.  Then we would like to group isolates based
> on the significant factors (e.g. if 'farm' is significant, then
> build a dendrogram grouped by farm).        In order to do this we
> must first determine the appropriate data structure that defines the
> 'genetic relatedness' variable.  We can't model a dendrogram
> statistically, but we could model the data that defines the overall
> dendrogram.  The matrix of correlations could be one data structure,
> but these numbers are not for individual isolates, but pairs of
> isolates.  The risk factors are for individual isolates (not pairs).
>  Can we define a relevant outcome(s) individual isolates?  Perhaps a
> measure
> (or series of measures) for each isolate that capture the data that
> defines the correlation matrix.  Perhaps a 'rooted' tree would serve
> this purpose so that all study isolates are compared to the same
> reference. There are probably other options as well.
>
> I hope this at least gives you an idea of our proposed
> approach/needs and can serve as starting point for our upcoming discussion
> Cheers,
> Dave
>
> David G. Renter, DVM, PhD
> Veterinary Epidemiologist
> Agri-Food Systems Branch
> Food Safety Division
> Alberta Agriculture, Food and Rural Development
> Phone: 780/422-0275
> Fax:  780/422-3438
> ------- End of Forwarded Message -------
>