Discrete models come up in virtually all application areas but one of them plays for us a particular role: molecular biology. Bioinformatics gives us a source of problems to study and, on the other hand, an excellent field to apply and to test our ideas and algorithms. We are mainly interested in fingerprints of biological phenomena in genomic or protein sequences. Those fingerprints are described in terms of patterns, and one of our main objectives is to identify, search and analyze those motifs using discrete algorithms and probabilistic analysis.
Below we briefly describe some of the bioinformatics projects we had working on.
We developed the YASS software for local similarity search in DNA sequences. YASS is a program to perform a local alignment of DNA sequences (fasta or plain format). YASS uses transition constrained seeds combined with statistical criteria to detect and extend similarity regions.
We have been working on the design and analysis of seeds for pattern matching and similarity search.
Some sites in the non-coding part of the genome are directly involved in the transcription regulation. The knowledge of those sites would allow us to identify co-regulated genes, to determine associated regulatory mechanisms and possibly to bring out proteins with unknown functions.
In the framework of the theme Bioinformatique et applications à la Génomique of the Pôle de Recherche Scientifique et Technologique (PRST) Intelligence Logicielle, we work on the identification and classification of regulatory sites in the Streptomyces coelicolor bacterium, in collaboration with scientists of the Laboratoire de Génétique et de Microbiologie de l Université Henri Poincaré de Nancy (Pierre Leblond, Bertrand Aigle). Note that this bacterium presents a particular interest, as more than seventy percent of the known antibiotics are produced using bacteria of the Streptomyces family.
Our ultimate goal consists in identifying binding sites of sigma-factors in promoter regions of the Streptomyces coelicolor genomic sequence.
Another project addresses the issue of repeated DNA sequences, occurring several (more than two) times in a genome. From the computer science viewpoint, this project is based on the YASS software, that can be used to identify pairs of repeated sequences in a given genome. On top of YASS, we created another software for computing clusters of repeated sequences. In order to exclude parologous genes, we restricted our search to intergenic regions only. The approach has been applied to identify clusters of repeated sequences in proteobacteria genomes, in particular in Neisseria meningitidis serotype A and B. Obtained clusters have then been analyzed using various bioinformatics techniques.
Here, our general goal was to study, using computer methods, the influence of the DNA curvature on the efficiency of binding sites of some proteins. In particular, we focused on the gene regulation possibly mediated by global regulatory proteins H-NS/FIS in bacteria.
The main difficulty of such kind of analysis is the absence or a degenerative form of conserved DNA-binding sequence motifs for both H-NS and FIS proteins. H-NS is described to bind non-specifically to DNA and prefers intrinsically curved regions. Based on this knowledge, we used CURVATURE software in order to predict possible H-NS binding sites upstream of rrn operons in proteobacteria containing H-NS. rrnB operon in Escherichia coli is known to be regulated by FIS/H-NS and regulatory region of this operon contain intrinsically curved region. We analyzed regulatory regions of other six rrn operons in this bacteria and found a curved region, with the center located approximately at -90 - 110 relative to the transcription start site. Analysis of regulatory regions of predicted rrn operons in other proteobacteria showed a high degree of DNA curvature upstream of transcription start site of rrn operons, although the position of the center of curvature in various bacteria differs. This work is a first step towards a general analysis of protein binding sites, using DNA curvature information.