Arkansas Study Targets Genetic Neighborhoods for Bacteria

University of Arkansas System Division of Agriculture

By John Lovett

University of Arkansas Division of Agriculture

FAYETTEVILLE, Ark. — When it comes to identifying harmful bacteria, it helps to look at the company their genes keep.

Researchers with the Arkansas Agricultural Experiment Station, the research arm of the University of Arkansas Division of Agriculture, used a machine-learning approach to study not just which genes a bacterium has, but also where those genes sit next to each other. That "genetic neighborhood" helped researchers distinguish disease-causing strains of Enterococcus cecorum, a poultry pathogen, from nonpathogenic strains.

While many strains of E. cecorum are harmless, others can cause arthritis, bone infections and lameness in poultry, which creates animal welfare concerns and economic losses.

"The machine-learning model recognizes patterns in gene order much like it recognizes patterns in language," said Aranyak Goswami, a computational biologist and assistant professor with the experiment station's Center for Agricultural Data Analytics. "Many studies examine genes individually. We are looking at their genomic locations, order and neighboring relationships."

The novel method was used to distinguish disease-causing strains of E. cecorum from nonpathogenic ones by analyzing how neighboring genes are organized within the bacterial genome. Goswami and his colleagues recently published the results of their proof-of-concept study in the journal Frontiers in Microbiology , showing that patterns in neighboring gene arrangement can help distinguish disease-causing strains of E. cecorum from harmless ones.

Goswami said the idea to study what the researchers call "genomic-island cassette architecture" originated with Rushikesh Lagad, a study co-author and doctoral student in the department of animal science. Goswami is also affiliated with the departments of animal science and poultry science in the Dale Bumpers College of Agricultural, Food and Life Sciences at the University of Arkansas.

"While reading the literature, I kept wondering whether we were missing part of the story by looking at genes one at a time," Lagad said. "That led me to study how genes are arranged together within genomic islands and whether those genetic neighborhoods could help distinguish pathogenic strains from harmless ones."

Neighborhood watch

Traditional genomic studies often focus on whether individual genes are present. While that approach is valuable, Goswami said, it often overlooks how neighboring genes are arranged and function together.

He compares the difference to looking at houses on a street. Rather than looking inside individual houses, the researchers examined the entire neighborhood to identify broader and more distinctive genomic structures.

Goswami and his colleagues examined this cassette architecture, which is the way groups of neighboring genes are organized, within larger sections of DNA called genomic islands. Unlike the rest of a bacterium's genome, genomic islands are often acquired from other bacteria and frequently carry genes that help bacteria survive, spread or cause disease, including antibiotic-resistance genes.

Within those islands, the researchers studied how groups of neighboring genes were organized to learn whether harmful strains had different neighborhood patterns than harmless strains.

Developing a map

Existing methods for monitoring E. cecorum often rely on culturing bacteria or screening for specific genes. Goswami said those approaches can miss broader patterns because they focus on individual genes rather than how neighboring genes are organized.

Instead of treating a genome like a shopping list, they treated it more like a map. The researchers examined where genes were located and which genes tended to occur nearby, not just whether a particular gene existed.

"Very few groups study this type of interaction architecture," Goswami said.

The researchers analyzed the genomes of 145 E. cecorum strains collected from poultry, including 95 nonpathogenic strains and 50 pathogenic strains capable of causing illness. The analysis found that disease-causing strains were more likely to contain genomic islands enriched with genes involved in antibiotic resistance and the movement of genetic material between bacteria.

A tool for future surveillance

Goswami emphasized that the method is intended as a research tool rather than a diagnostic tool.

Although the approach showed strong performance within the study, additional validation will be needed before it can be used routinely to monitor poultry flocks or identify emerging disease-causing strains, he said.

Still, the findings suggest that the organization of genes, not simply their presence, may provide valuable clues about how bacterial pathogens evolve and emerge.

"This gives us another way to understand these bacteria," Goswami said. "If we can recognize patterns that distinguish harmful strains from harmless ones, that could eventually help improve surveillance and generate new hypotheses about how these pathogens evolve."

The study also shows how computational biology and machine learning can complement traditional microbiology by revealing meaningful patterns hidden within large genomic datasets.

The study was co-authored by Shakil Rafi, a mathematician who worked on the project as a postdoctoral research fellow in data science and bioinformatics before joining the University of Arkansas College of Engineering as a teaching assistant professor.

Multiple-species pipeline

Although the study focused on a poultry pathogen, Goswami said the computational pipeline could be adapted to many other bacteria. He pointed to collaborations with researchers studying bee pathogens as one example and said the same approach could eventually be applied to bacteria affecting humans, livestock, wildlife and plants, provided enough genomic data are available.

"The approach is not limited to one bacterial species or habitat," Goswami said. "If sufficient genomic information is available, we can adapt this pipeline."

Goswami said the team is beginning to adapt the computational pipeline to study Enterococcus faecalis, a close relative of the poultry bacterium that is a common cause of hospital-acquired infections in humans. The researchers also plan to apply the approach to E. coli and other bacteria to better understand how nonpathogenic strains evolve into disease-causing pathogens.

"Our goal is not just to publish a paper," Goswami said. "We wanted to build a pipeline that researchers studying any bacterial pathogen can use."

The research began when poultry breeding company Cobb-Vantress approached Goswami for his expertise in machine learning and genomics to study E. cecorum. Seed funding from the Arkansas Research Alliance helped launch the collaboration.

The study included publicly available bacterial isolates from Arkansas, along with genomes from other sources, allowing the researchers to evaluate the approach using strains associated with the state's poultry industry.

To learn more about ag and food research in Arkansas, visit aaes.uada.edu

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.