
\documentclass[12pt,a4paper]{article}

\usepackage{graphicx}
\usepackage[sort]{natbib}
\usepackage[backref]{hyperref}
\usepackage{setspace}
\usepackage{anysize}
\usepackage{subfigure}
\setcounter{tocdepth}{4} \pagestyle{headings} \doublespacing

\marginsize{4cm}{2cm}{2cm}{2cm}

\renewcommand{\baselinestretch}{1.6}

\begin{document}

\author{Emily Jefferson}
\title{Manual For PIP (Protein Interaction Predictor)}
\maketitle

\section{Outline of Program}
The PIP package is a collection of classes which can be used to analyse the interacting domains seen in structural data. 

\section{How to create the System}

\subsection{Precursors to generating the System}

The following steps need to be taken in order for the system as a
whole to work. However, some of the modules within the system do not
require all of the following to be implemented. Therefore, the
description of each module (next section) states which steps are
necessary.

\begin{description}

\item [Insure that you enough disk space]
{\em 60GB of available space for permanent use and another 40GB for
extra temporary files generated}. The database takes up about 40GB
of space. However, there are also some other files which need to be
generated to have a working system. These include PDB files each
assembly whole assemblies generated with unique numbering from the
MSD, whatIf files, pdb files for each domain structure, Jess
templates and Jess results. These together require another 60GB.
Therefore, it is recommended that you have available 100GB of space.
All of this space is not required once the system is up and working
as some of the results are temporary and once the results have been
included in the database they can be deleted. When the temporary
files can be deleted is described below.

\item[Install Java 5 or later]
The code is written using Java 5 language specificities and so a
version of Java 5 is required to run this system. Java 5 can be down
loaded at http://java.sun.com/j2se/1.5.0/download.jsp

\item[Connection to the MSD Database]
A connection to the MSD database is required.

\item[Install Fast Objects]
In order to create the object orientated database from the MSD
database (as described above) Fast Objects first needs to be
installed. If you want to use another implementation of JDO see the
section below. The Fast Objects web site has all the information
about installing the program. You just need to get an academic
licence (community edition) and then run the installer. From there
all the other setup things are done with in my code.

\item[Install Ant]
Ant is a platform independent method for making programs. Apache Ant is a Java-based build tool. In theory, it is kind of like Make, but without Make's wrinkles. Instead of a model where it is extended with shell-based commands, Ant is extended using Java classes. Instead of writing shell commands, the configuration files are XML-based, calling out a target tree where various tasks get executed. Each task is run by an object that implements a particular Task interface. Ant is
used to generate the code for the system. Ant can be found at http://ant.apache.org/.

\item[Install Blast]
Blast is need in the prediction of protein-protein interactions.

\item[Install Pymol]
Pymol is needed to view interacting domains if using the interaction predictor part of the program. Pymol can be downloaded from http://pymol.sourceforge.net/.

\item [Install Scanps]
As described above the program Scanps is required for one the steps
in the system. Scanps is available under academic licence from
www.compbio.dundee.ac.uk. However, it is also packaged with this
program and so can be installed from the scanps dir. Explain.

\item [Programs required for Template Searching]
The template functionality can be turned off and so if this
functionally is not required the following steps do not need to be
completed.

\begin{description}
\item[Install Jess] Jess is needed to search for interacting site templates.

\item [Install HMMER]
HMMER is required for the generation of templates of interacting surfaces. HMMER can be found at http://hmmer.wustl.edu/.

\end{description}

\item[Set the properties of the System]
Where each of the programs is installed and were you want to put
each of you results is system independent therefore these properties
need to be set by the user.

These properties are set in the properties.txt file. A description of each property is shown below. The properties.txt file contained within the PIP package contains examples of the choices.

DATADIR = this is the directory you want to put the data from the analysis of the database
WHATIFDIR = this the directory where you want to put the contact information generated by whatif 
RESSUBDIR = this is the directory where you want to store the pair potentials of each residue pair
BLASTDIR = this is the directory which contains the Blast program
PDBFILES = this is the directory where you want to store all the generated PDB files
TEMPLATEDIR = this is the directory where you want to store all of the generated templates
WEBBROWSER =  this is the path to the webbrowser that you want to use to view the results
PYMOL = this is the path to the pymol program
MSDDRIVER = this is the oracle driver for the MSD eg. oracle.jdbc.driver.OracleDriver
MSDURL  = this is the jdbc dirver e.g. jdbc:oracle:thin:@hornet:1521:MSDSD
MSDUSER = this is the username to access the msd e.g WHOUSE1
MSDPASSWD = this is the password to access the msd e.g.abc123
javax.jdo.PersistenceManagerFactoryClass=com.poet.jdo.PersistenceManagerFactories
javax.jdo.option.ConnectionURL= this is where the database is stored e.g. fastobjects://LOCAL//grid/zeus2/emily/DATABASE.j1

Not all of these properties need to be set from the beginning as some are only required to perform various sections of the PIP package. Therefore, in the description of each section the properties which need to be set are described.

\end{description}

\subsection{Generating the System}

For minumum requirements of the basic system to work are as follows:

\begin{itemize}
\item
Install fastobjects community edition
\item
Install Java 5 (or later version)
\item
Install Ant
\item
Connection to MSD
\item
40GB of space for object orientated database
\item
The following directories in the properties.txt file will need to be set
DATADIR
WHATIFDIR
PDBFILES
MSDDRIVER
MSDURL
MSDUSER
MSDPASWD
javax.jdo.PersistenceManagerFactoryClass
javax.jdo.option.ConnectionURL

\end{itemize}

The generation of the system requires 3 main steps: generation of the database sturcture, generation of the PDB files and WhatIf Contact files and creation of the domain interacting pairs. The generation of the database structure and the generation of the PDB files and WhatIf files are not dependant on each other and can be run in parallel. Both the generation of the database structure and the generation of the PDB files and WhatIf files steps are required for the third step of creation of the domain interacting pairs.

\subsubsection{Generation of the Database Structure}

Maybe discuss the RII function if applicable to cross platform.

\begin{description}
{\em Ant Task:} 
Use "make" ant task to build the code present in the PIP directory. This produces several errors to do with the enhancing of standard Java libary classes. These errors are not a problem because the classes they are enhancing do not contain any data. However, this will mean that ant can not be used to run the database and so once the ant task has been used to enhance the classes the main files need to be called from the command line or from in your IDE (without compiling first as the ant task has already done this).
{\em Create JDO database instance} 
To create the database from scratch go to the place where you want to store the database and run the following command:-
	ptj -ep classes -register -xc -database DATABASE.j1
This creates a database called DATABASE.j1.
{\em Populate JDO database structure} 
To populate the database with the structure of the object you wish to persist the following main class must be then called
	EII.CreateDatabase.CreateNewDatabaseInstance
{\em Populate JDO with Entries} 
To populate the database with the Entry and Domain Information from the MSD the following main class must be then called

	EII.CreateDatabase.PersistEntries
	
This will take a long time to run as it connects to the MSD with one connection and then adds each generated object to the new JDO Database. An alternative is to run a cluster job instead. This job connects to the MSD in parallel and stores each object as a serialised object. Each one of these serialised objects can then be added to the JDO database. This is much faster. To do this each node requires a list of EntryIds form which to querry the MSD. Firstly create a directory called "entryIds" in the directory DATADIR that has been set in the properties.txt file. Lists of EntryIds are created into this directory using the following command

	java EII.CreatePeripheralData.EntryIdListCreator 100
	
The number 100 is passed in as a variable to determine how many EntryIds there should be per File. The greater the number the more Entries will be processed but one cluster node. 

Now make a directory called msd_serialised in the DATADIR directory. The cluster job can then be run using a perl script called "runBatchMSDSerialise.pl" which is present in the PIP directory. To run this script on the cluster firstly count how many files there are in the entryIds directory using (from within the entryIDs directory)

	ls |wc-l

Using the number returned (lets call it foo) the cluster job can then be run. Firstly change the paths in the perl script runBatchMSDSerialise.pl to suit your system. Then from within the PIP directory run the cluster job as follows

	qsub -N CreateSerialisedObjects -cwd -t 1-foo:1 runBatchMSDSerialise.pl	
	
This should take about ..... hours to run. Once this has finished running delete all of the error logs in the PIP directory using 

	find -name "CreateSerialisedObjects.*" | xargs rm

To then read in the serialised objects to the JDO Database run

	java EII.CreateDatabase.ReadInSerialisedObjects
	
{\em Populate JDO with Classified Domains}	
The database now contains a list of all entries in the MSD and the relevant information relating to each entry.
As most of the analysis of the database is based on domains the domains in the database are classified depending on there domain class in to an easy to use structure. This structure is generated using 

	java EII.CreateDatabase.GenerateSetsOfDomains

\end{description}	

\subsubsection{Generation of the PDB Files and WhatIf Contact Files}

In order to generate the domain interactions present in the database a set of Whatif generated interaction data is required. To generate the WhatIf files a set of PDB Files for each Assembly needs to be generated. These PDB
files use a numbering system which is different from the PDB files
that can be downloaded from the PDB group as the MSD numbering is
used. These files only contain the ATOM lines and no other
description.
		
\begin{description}
{\em Create PDB files} 
To generate each PDB file run

	EII.CreatePeripheralData.GeneratePDBFilesForEachAssembly	

The problem with this is again it is slow as it uses only one connection to the MSD. To run this on the cluster instead first create a directory called "assemblyIds" in DATADIR. Then run 

	EII.CreatePeripheralData.AssemblyIdListCreator 100

The number 100 is passed in as a variable to determine how many AssemblyIds there should be per File. The greater the number the more Assemblies will be processed but one cluster node. 

The cluster job can then be run using a perl script called "runBatchPDBFiles.pl" which is present in the PIP directory. To run this script on the cluster firstly count how many files there are in the assemblyIds directory using (from within the assemblyIDs directory) 

	ls |wc-l
	
Using the number returned (lets call it foo) the cluster job can then be run. Firstly change the paths in the perl script runBatchPDBFiles.pl to suit your system. Then from within the PIP directory run the cluster job as follows

	qsub -N CreatePDBFiles -cwd -t 1-foo:1 runBatchPDBFiles.pl
	
This should take about ..... hours to run. Once this has finished running delete all of the error logs in the PIP directory using 

	find -name "CreatePDBFiles.*" | xargs rm
	
{\em Create WhatIf Contact Files}

The Whatif Contact results files can then be created using runWhatif.pl within the PIP directory. Firstly the paths will need to be changed to suit your system and then run

	perl runWhatif.pl
	
Again there is a cluster job for this as well - but it does not work yet!

\end{desciption}

\subsubsection{Creation of the Domain Interacting Pairs}
As stated above this step requires that the above two steps are completed.

\begin{desciption}
{\em Populate the JDO Database with DomainInteractions}

\end{desciption}

\subsection{Analysis of the Data}

The minumum requirements for the analysis of the database are as follows:

\begin{itemize}
\item
All of the minmum requirements of the generation of the database structure step.
\item
\item
\item
\item
\item
The following directories in the properties.txt file will need to be set

\end{itemize}

\subsection{Protein Interaction Predictor}

\subsubsection{Using the Protein Interaction Predictor}

\begin{desciption}

\end{description}

\subsubsection{Testing the Protein Interaction Predictor}

\subsection{Interaction Site Templates}

\subsubsection{Generating Interaction Site Templates}

\begin{desciption}

\end{description}

\subsubsection{Testing Template Searching}

\subsubsection{Using Templates to Predict Interactions}

\end{document}
