Logo
About usInnovation CMChallengesEuropa2iEntrepreneurshipR&D&I SearchAgentsEventsReports
en
METHOD FOR TARGET METAGENOMIC SEQUENCING OF MULTIPLE MICROBIAL GENE FAMILIES WITH FULL PHYLOGENETIC COVERAGE IN A SINGLE EXPERIMENTCM Patents

Índice de la ficha

Updated at
24/07/2026
Numero publicacion
EP.4730339.A1
Fecha publicacion
22/04/2026
Numero solicitud
EP20240383141
Fecha presentacion
17/10/2024

En detalle

Resumen

The vast majority of prokaryotic microorganisms remains uncultured and are not abundant enough to be effectively studied using metagenomic sequencing methods. Here, we present a novel PCR-free target capture sequencing approach that enables targeting specific genes or functions with full phylogenetic coverage in a single sequencing experiment, proving also a higher sensitivity and precision than regular shotgun and amplicon-based approaches.

Reivindicaciones

1. An in vitro method for providing a collection of sequencing probes that enable targeting, preferably in a single sequencing experiment, one or various gene families or functions with full phylogenetic coverage and high sensitivity, said method comprising the steps of: a. Step 1. Gene family selection: Compiling a list of marker gene families or functional orthologous groups (OG) by, preferably, identifying genes from a database within or associated to a specified group of molecular functions. b. Step 2 Data mining: Identifying and providing from one or more databases, a data corpus or dataset of nucleotide sequences having a percentage of identity with the genes and OGs identified in step 1 of at least 60%, 70%, 80%, 90%, 95%, 99% or more; c. Step 3. Design of capture-kits: Analyzing the data corpus or dataset provided in Step 2 in order to reduce its dimensionality to make it accessible for a single targeted sequencing experiment, preferably while maintaining the same phylogenetic coverage and sensibility. d. Providing the design of the panel or collection of sequencing probes that enable targeting, preferably in a single sequencing experiment, one or various gene families or functions selected in Step 1. e. Optionally, manufacturing or producing the panel of step d. 2. The method according to claim 1, wherein in Step 2 the data corpus of nucleotide sequences is provided by retrieving for each of the selected OGs of step 1, the sequences from (i) global genomic repositories and (ii) public metagenomic datasets. 3. The method according to claim 2, wherein Step 3 is carried out by a method comprising the following steps i) to iv) for each targeted OG (tOG) included in a kit: i. Decomposing into non-overlapping chunks of from 10 - 200 pair bases, preferably from 50 to 100 pair bases, more preferably about 80 pair bases, all or part of the nucleotide sequences comprising the dataset obtained in Step 2 for each of the tOG, herein referred to as probes. ii. Aligning the nucleotide sequences comprising the dataset obtained in step 2 into a multiple sequence alignment (MSA), where the conservation level of each column is calculated. iii. The multiple sequence alignment produced is used to generate a phylogenetic tree of all or part of the non-redundant sequences. iv. All or part of the generated probes in step i) are compared in an all-against-all fashion to calculate their pairwise sequence similarity, computed as the % of nucleotides identities observed between two probes and provide an all-against-all probe comparison matrix. wherein the method further comprises the following steps: v. Using the all-against-all probe comparison matrix of step iv to generate clusters of probes at the NSI threshold, wherein the NSI threshold is selected from one or more values in the range of from 30% to 100%. vi. Mapping all or part of the probe clusters of step v back to the same original genomic and metagenomic databases used in Step 2 using BLASTN searches. vii. Analyzing to compute the ratio of positive matches (i.e. hits against a gene belonging to the expected OG) over negative matches (i.e. hits against a gene belonging to an unwanted OG) of the hits produced in step vi. viii. Discarding all probe clusters with a ratio of positive matches lower than 1.0 (i.e. probes producing at least one mismatch) from the set. ix. Mapping all or part of the remaining selected probe clusters to the MSA generated in step ii), preferably by recording the exact position of matching residues of the original probe on the MSA. x. Discarding probes matching columns of the alignment produced in ii with a residue conservation score, preferably of below 0.75. xi. Mapping back the final set of selected probes to the original sequences selected in Step 2, preferably by using BLASTN, and preferably calculating: number of original sequences recovered by the probes (RecovT), the average length of the original sequence covered by the probe-selection (RecovLen) and the phylogenetic coveraged based on the the generated in substep C (PhyloC) 4. The method according to claim 3, wherein for each tOG the total number of probes selected (Nprobes), together with their RecovT, RecovL and PhyloC values obtained for each clustering experiment under each NSI are automatically evaluated to select the set of probes that maximizes RecovT, RecovL and PhyloC, while minimizing NProbes. 5. The method according to claim 4, wherein the set of probes that maximizes RecovT, RecovL and PhyloC, while minimizing Nprobes, for each tOG, is combined into a single dataset, and providing the design of the panel. 6. The method according to claim 5, wherein the method further comprises the manufacturing or production of the panel of probes. 7. A panel obtained or manufactured or produced by the method of claim 6, wherein the probes are polynucleotides, preferably DNA or RNA, probes. 8. Use of the panel according to claim 6, for the capture and target enrichment assays directly on environmental DNA samples, prior to shotgun sequencing. 9. A kit comprising the panel obtained or manufactured or produced by the method of claim 6

Etiquetas

Inventores
Huerta Cepas JaimeGonzalez Bodí SaraPérez Cantalapiedra CarlosSanchis López ClaudiaPokorny Montero Cristina Isabel
Solicitantes
Consejo Superior de Investigaciones CientíficasUniversidad Politécnica de Madrid
Clasificacion ipc
G16B 25/ 20 A IG16B 30/ 10 A IG16B 40/ 00 A I
Logo

Innovation CM
Challenges
Europa2i
Entrepreneurship
R&D&I Search
Agents
Events
Reports
About us
Contact
Give us your opinion
Cookies
Legal notice
Privacy

© Copyright Espacio Madrileño de Investigación e Innovación 2026