Thursday, September 30, 2010

Notes from ETNA Summer School 2010






Prof. Atanas Atanassov from JGC giving a speech

This year ETNA Summer School is jointly organized by Joint Genomic Centre (JGC) and AgroBioInstitute (ABI) which are located at the Department of Biology in Sofia University. After a tour around the facilities, we were given an interesting introduction about the research interests in Bulgaria. Besides rose oil and grapewine(and Black Sea!), this country has a lot more economically important areas. ABI also focused on important crop such as wheat, barley, berries, lactic acid bacteria, thermophile bacteria, honey bee, medicinal herbs, and animal breeding.

The first day is followed by 10 presentations. Most plant-specific talks focused on grapevine genomics and breeding. A plant totally unfamiliar to me. Grapevine or vitis vinifera is grown for wine and table grapes. Wine yard covers 59% of Europe. It’s a perennial crop, grafted and high in genetic diversity, Quality of wine is greatly influenced by environment . The sugar content in the grape is important for fermentation. Although the genome sequences are available, many genes are still unknown and unannotated. Many are unique to the grapevine. It's a highly heterozygous plant. Difficult transformation. It’s an ancient allopolyploid before going through diploidization 2(6+6+7)=38. However, the American grapewine (such as Muscadinia) contains 40 chromosomes. The studies focused on disease resistant, cold or drought tolerant, budding responses and berry development using a combination of transcriptomic (mainly using microarray) and metabolomics.


Visiting the wine yard in Starosel

Metabolomic is an important component of this course. Something new to me. The participants were introduced to basic concepts of metabolomics, sample preparation, detection using GC-MS and most importantly analysing metabolomic data using bioinformatics tools. Every metabolite profile is closely associated to phenotype. For example, different developmental stages of grapewine berry has different concentration of compounds. Metobolomic helps scientists to link the change of metabolites to phenotype such as taste or colour. It's interesting how metabolomic can be integrated with transcriptomic data to produce more biological significant results.


Practical sessions

The most interesting talk is epigenetics in plant breeding presented by Prof. Atanasios Tsaftaris. First, he explained the methylation and histone modification mechanisms and some classical examples of plant epigenetics. Epi- means "above" so epigenetics means "above genetics". It's genetics that doesn't follow Mendel's law. I also learn Genetics Imprinting - Expression of only one allele from the parents due to suppression of the other allele caused by methylation. It's now known that epigenetics play an important role in sensing the environment, control of flowering time and seed development. It is also the cause of somaclonal variation in plant tissue culture and why clones in the field don't perform the same. Two years ago, a Nature paper about Arabidopsis epigenome was published but one can't truly appreciate that paper until he/she understand epigenetics and its implication in plant biology. Read this paper!

This summer school has provided me great opportunity to interact with researchers and students. The topic about funding problem was brought up during coffee break. Research funding has been reduced due to economic crisis and I believed it happens everywhere. The grant application criteria in every country are different. In Europe, there's national and EU funding. A national funding is supported by the government. In the Netherlands, the project must be supported by a few private companies before getting grant approval from government. To secure a EU funding, the project must involve two or more countries with a common research interest. In Malaysia, almost all research funding came from the government. Compared to Bulgaria, Malaysia government or universities have been very supportive to postgrad students by providing scholarship and tuition fees waiver. Now I finally have to agree that we have lots of funding and opportunities in Malaysia. It's up to the Malaysian researchers' initiative and creativity to make use of the available resources. So this is one BIG take home message that I wanna tell my colleagues.

Read more...

Back from ETNA Summer School



I was away to attend the 3rd ETNA Summer School in Sofia, Bulgaria. My 10 days stay in Sofia has been very pleasant. Besides learning a lot of new things, I made lots of new friends. :)

The ETNA summer school is sponsored by European Union to provide training and networking opportunity to young researchers. The application is open for PhD students and post-doc early in their career. I stumbled upon the website last month and I thought I will give it a try. This year focus is plant genomics and bioinformatics in plant breeding. The school invited speakers from all over Europe and organized practicals for the the "omic" technologies. Travel and accommodation is sponsored. The next summer school will focus on system biology. Application is open in April next year. So keep your eyes open!

Read more...

Wednesday, August 11, 2010

Illumina Seminar on developing MAS on agrigenomics in plant



I just came back from a seminar on plant Marker Assisted Breeding organized by Illumina yesterday. Since I'm waiting for a 8-hours script to complete (blame my bad programming skills), I will post something about the seminar. Last month, Illumina has announced to give away 10G of sequences to any Malaysian scientist who come up with the best 5000-words proposal. The deadline is 31 Aug 2010.

Back to the topic. The speaker is Dr. Richard Hodgson from Illumina US. He has vast experience in developing disease diagnostic methods for agriculture and aquaculture. He got involved in many breeding projects such as chili, coconut, shrimps and now he has a liking in durian. He introduced a relatively new approach called Genomic Selection(GS).

Unlike Marker Assisted Selection(MAS), GS is based solely on genotyping and estimation of breeding values. It has the advantage of capturing small gene effects not detected by QTL mapping. First, the breeders must have a large training population with known genotypes and phenotypes. It takes advantage of the cheap genotyping to screen tens to hundreds of thousands SNPs markers for large number of seeds/seedling. The number of markers depends on the diversity of the populations. The more diverse, the more markers needed. The selection is based on how closely the genotype matches the training population and breeding values is estimated based on predicted phenotype. Without phenotyping, the breeding and selection cycle is reduced significantly.

Here is a good introductory article about GS here. "In simulations, the correlation between the true breeding value of unphenotyped experimental lines and that predicted by genomic selection has reached 0.85. Genomic selection accuracies depend on a trait’s underlying genetic architecture, the level of linkage disequilibrium in the crop population relative to the marker density available, and the statistical methods used." Another paper mentioned why GS is better than association mapping.

Studies in maize and wheat has demonstrated success using GS. The question is how well does it work? How to apply it in non-model plants? Does it require a linkage map and location of each markers on the chromosomes? One thing for sure, you need to assemble a consortium, sequence a lot of varieties/lines and design a SNP chip for this purpose. Here's where Illumina play an important role in providing technologies bla bla bla... zzz.

Read more...

Monday, July 12, 2010

Bioinformatics lecture and course: August 2010





Lecture on Systems Biology and Data Integration
Title: Supporting
Systems Biology and Data Integration
Speaker: Prof. Chris Rawlings
Date: 3 August 2010
Time: 2-5pm
Venue: Room 304 & 305, KL convention Centre

Bioinformatics Advanced course
Speaker : Prof. S. Halgamuge, University of Melbourne
Date: 27- 29 September 2010
Venue: Ritz Calton Hotel, KL
Organizer: Ironix-Continuing Education
Registration fees: 980 euro (early bird); 1,280 euro (late)
For more information, click here.

Read more...

Thursday, June 24, 2010

Can't believe half year gone!



July 2010 marks the beginning of a new semester. Time flies. Lately, I noticed two questions that I was regularly asked: 1) How's your project going?; 2) When do you finish your PhD? *Sigh. Things that I should have done 3 months ago is still hanging in my list. Looking back, it's not a bad first half of the year!

A few days ago, a new desktop arrived in my lab. Finally! I had to install it twice because I couldn't partition the hard disk correctly. There is no such thing as root and swap space the last time I remember! Then, I spent two days updating the files, fixing the screen resolution and installing software. Problems never fail to arise when I use Linux. Next, I'm planning a Linux introduction workshop for my labmates.

Today it's the last day of Kolokium FST. It's an annual event where the final year postgrad student present their work. Coincidently, there is a seminar on marine biology and also one held in Inbiosis. Less people attended it compared to last year. This is a great opportunity to shamelessly helped myself to the food.

Read more...

Friday, June 18, 2010

Preprocessing of NGS reads: Trimming and filtering



The last few MGRC seminars I attended has been emphasizing the importance of pre-processing of NGS reads. SNPs detection relies heavily on read quality. Softwares like Maq and Samtools use quality score. But it's a different story when it comes to assembly because assembly doesn't use quality score. Some think that trimming is only necessary in very high coverage data.

Only the past few months, more tools on read quality assessment are available for the wider community. A recent IlluminaGAII 1.3+ Pipeline documentation reported that a run of bases with a quality score 2 or symbol 'B' is unreliable and should be removed (View discussion here). With increasing understanding and the tools available, I think the community can make a better decision to trim or not to trim.

Here is a list of trimming tools/scripts I found online:
1. HT Sequence Analysis with R and bioconductor
2. HTSeq by Simon Anders.
3. Softtrim.R by Jeremy Leipzig
4. TrimBWAstyle.pl by Joe Fass
5. Solexa_Sig2
6. Biopython

Bioconductor ShortRead package has been available since 2008. It requires a bit of R knowledge. It's only recently I realized that it has a function to trim reads to desired length. It only can perform right trimming.

HTSeq is able to trim adaptors but not removing low quality reads. The ht-seq-qa script is particularly useful. It generate a nice plot to show distribution of quality score over read position.

Here's how I perform trimming:
First, I installed Bioconductor ShortRead package. Make sure you use the latest version of R. Please note that trimming using R is RAM intensive. In my case, it took an hour to process a 1Gb file using 12Gb RAM computer. The program sometimes get killed or return an error message saying "Error: cannot allocate vector of size 1.5 Gb".

Then, convert the quality score of your fastq file to Sanger Phred score using IllQ2SanQ.pl script from UC Davis. Next, I used Leipzig's Softtrim R script. I can choose the min quality score, min read length of trimmed reads and enable left trimming. This script works better for me. Click here to see how to use his script. Lastly, I separated the trimmed reads into paired and single reads using a Python script.

Here's how my trimmed reads look like:

Before trimming

After trimming

*Plots are generated using htseq-qa script by Simon Anders

Read more...

Thursday, May 20, 2010

Linux for bioinformatics training



I don't think anyone should learn bioinformatics without learning how to use Linux OS. Most of us are self-taught. Won't it be great if training course is easily available. Many times we are not aware. Only recently I was informed by a colleague that such courses can be organized in the computer faculty. And I just missed one! Imagine my frustration. Until I received an email this morning.

Workshop on Introduction to Linux for Bioinformatics
Date: 25-26 May 2010
Venue: Postgraduate Laboratory, Bioinformatics Department, Science Faculty.
Organizer: CRYSTAL, UM
Registration fees: RM500 (students); Rm700 (Academia)
For more information, email crystal_seminar@um.edu.my

Read more...

Thursday, May 13, 2010

May 2010: NGS Seminar & lecture



There are two important Next Generation Sequencing events this month:

MGRC Lecture
Title: Application of Next-Generation Sequencing Technologies as Shared Research Resources
Date : 25 May 2010
Time : 2 - 5 pm
Venue : Plenary Theatre, Level 3, KL Convention Centre

Illumina South Asia Pacific Seminar
Date: 31 May 2010
Time: 8.45am
Venue: Cempaka Room, Hotel Equatorial, Bangi
Organizer: Sciencevision

Read more...

Thursday, April 8, 2010

Bioinformatics for Biologist: Installing standalone BLAST+ on linux



I'm sure most of us use NCBI blast on a daily basis. I use blast2sequences frequently. Sometimes I get frustrated because I cannot blast two set of sequences to each other. This leads me to explore options on how to run a local blast.

So what is standalone BLAST? The answer below is quoted from NCBI faq section:

"The StandAlone WWW BLAST Server allows you to set up your own in-house version of the NCBI BLAST Web pages. This can be accessed through web browsers on intranet web servers. You can set up the program to search your own custom databases or downloaded copies of the NCBI databases. The StandAlone WWW BLAST Server is available by anonymous FTP at ftp://ftp.ncbi.nih.gov/blast/server/current_release/."

Free software such as Bio-edit can perform local blast. Unfortunately, it cannot handle a huge amount of data and it only run on Windows. Another alternative is BLAT. Some people align short reads to genome using BLAT due to its speed.

The advantage of running a standalone BLAST is you can use the blast algorithm to search for your queries against the database you created (aka local blast). Your database can be a nucleotide or peptide FASTA file of your data or any data downloaded online.


The latest edition is BLAST+, an improved version of BLAST. The tar file can be downloaded from ftp://ftp.ncbi.nih.gov/blast/executables/. After untar, you will get the ncbi-blast folder with two folders inside: bin and doc. The commands:

>cd ncbi-blast-2.2.23+/bin/
#cd to where the blast executables are
>./makeblastdb -in database1.fa -dbtype nucl -out database1
#make the local database. Three files with extension .nrh, .nin and .nsq will be produced.
>./blastn -help # for more options
>./blastn -task blastn -db database1 -query query1.fa -out results1.txt -evalue 1E-50 -outfmt 6
#run blastn with blastn algorithm. Just type in database name, query file, output name and you can even select the E-value. Output format 6 presents results in table form.


Done!

Read more...

Wednesday, March 24, 2010

Notes from IUFRO Kuala Lumpur 2010




Bukit Melawati lighthouse, Kuala Selangor (In-conference tour)

I'm back from IUFRO Kuala Lumpur 2010 conference. It's a blast thanks to the committee members for their hard work (including myself :p). Just wanna post some short notes I gathered.

The conference opened with a keynote by Dato Freezailah who is the chairman of Malaysia Timber Certification Council. One of the interesting topics during the first day is about timber tracking. On the second day, Prof. David Neale presented a paper on adaptive and conservation genetics. I was truly captivated by his slides on the history of forest genomic approaches the past 20-30 years. I wasn't even born, imagine that!

On the 3rd day, tree genomics and bioinformatics workshop was held. Prof. Carl Douglas gave us a wonderful start on Popular genomics. It's amazing how many participants showed up for that session. The participants were eager to learn about Next Generation Sequencing and how to apply them in genomics, adaptive genetics and conservation. I presented during the workshop and got some good feedback from the audience. :-) Well, I could have done better.

The next IUFRO conference will be held in Florence, Italy and scheduled to be August 2011. The following conference will be in Kyoto, 2012. It's gonna be exciting!

Read more...

  © Free Blogger Templates Spain by Ourblogtemplates.com 2008

Back to TOP