- Job Title
- Postdoctoral Research Scientist (Bioinformatics) – Protist Genomics
- Post Number
- 1006194
- Closing Date
- 1 Oct 2026
- Grade
- SC6
- Starting Salary
- Salary: £39,000 - £46,500
- Hours per week
- 18.5
- Project Title
- Gordon and Betty Moore Foundation – Protist Omics at Scale
- Expected/Ideal Start Date
- 02 Nov 2026
- Months Duration
- 36
Job Description
Main Purpose of the Job
The Earlham Institute has been awarded funding by the Gordon and Betty Moore Foundation to develop Protist Omics at Scale, a three-year international methods-development programme.
the programme is run in partnership with the Scottish Association for Marine Science, home of the Culture Collection of Algae and Protozoa, and the Bigelow Laboratory for Ocean Sciences.
Protists represent the overwhelming majority of eukaryotic diversity, yet they remain hugely under-represented in reference genome databases. Their genomes are large, repetitive, frequently aneuploid and often recovered from mixed or low-biomass cultures, and the resulting data defeat standard assembly and annotation approaches. This project will systematically identify, diagnose and resolve those bottlenecks across three aims: cultured protists with substantial biomass, protists that grow only to low concentrations ,and single-cell genomes and transcriptomes of uncultured protists from the environment.
We are looking for a Postdoctoral Research Scientist / Computational Biologist to lead the computational core of the project: quality control, assembly, decontamination and co-biont separation, and structural and functional annotation across all three aims. Where existing tools fail, the postholder will diagnose why and develop what replaces them.
A defining requirement of the award is that the resulting methods are usable by the wider research community, so every workflow will be packaged, documented and released openly, with all data deposited under rich, FAIR-compliant metadata.
The post offers substantial scope for independent research. This is a rare chance to work at the interface of long-read sequencing, single-cell genomics, and eukaryotic genome assembly, on organisms whose genomes are largely unexplored.
The post is based at the Earlham Institute on the Norwich Research Park.
Key Relationships
Internal: Reporting to the Director and Principal Investigator, Prof Neil Hall. The post holder will work closely with the project Co-PIs (Dr Iain Macaulay, Dr Karim Gharbi), the project manager, the postdoctoral scientists and senior research assistants delivering the wet-lab components of the project, and with colleagues across the Institute's Core Bioinformatics, Scientific Computing and Data Infrastructure teams. They will also interact with the Research Faculty Office, the Communications team and the Institute's training group.
External: Co-PIs and researchers at the Scottish Association for Marine Science / CCAP and the Bigelow Laboratory for Ocean Sciences (Single Cell Genomics Center); the Gordon and Betty Moore Foundation and the project's external advisory board; the Darwin Tree of Life Genome Engine and the Wellcome Sanger Institute; the British Society for Protozoology and the wider international protist research community; WorkflowHub, Galaxy, protocols.io, ENA and NCBI; and collaborators and users attending project workshops, hackathons and training events.
Main Activities & Responsibilities
- Percentage
- Genome and transcriptome analysis
• Generate de novo genome assemblies for cultured protists, including chromosome-resolved assemblies using long-read assemblers integrated with other long-range information.
• Develop and apply bespoke assembly strategies tolerant of low-inputs, amplified material, and of single-cell genomes, including coloured de Bruijn graph approaches and the adaptation of existing assemblers to single-cell data.
• Implement read pooling and co-assembly strategies across multiple amplified cells of the same population to improve genome recovery.
• Assemble and evaluate short-read and long-read transcriptomes, supporting the analysis of isoform diversity, UTRs, trans-splicing, non-canonical splice sites and alternative polyadenylation in protist transcriptomes. - 25
- Decontamination, co-biont separation and chimera removal
• Remove and/or separate co-biont and contaminant genomes during assembly clean-up using the kmer-ord workflow developed at SAMS and established binning tools (e.g. Eukfinder, CONCOCT, SemiBin2), together with tools under development at EI, adapting them where necessary for single-cell and low-input data.
• Detect and remove physical chimeras introduced during whole-genome amplification using reference-based and de novo approaches, and benchmark read-length thresholds and protocol-level strategies that minimise chimera formation and data loss.
• Manually assess and curate assemblies to resolve mis-assemblies and to separate organellar genomes. - 15
- Genome and transcriptome annotation
• Deliver structural and functional annotation of assembled genomes and transcriptomes, integrating RNA-seq evidence where available..
• Assess assembly and annotation quality using BUSCO, k-mer completeness, contiguity and reference-based benchmarking, and contribute to the project’s community-aligned definitions of assembly contiguity and completeness. - 15
- Sequencing data quality control and triage
• Work with lab teams to rapidly analyse data for technical development work Perform QC of DNA and RNA sequencing libraries and of raw short-read, long-read (PacBio and ONT) and long-range (Hi-C, Illumina Constellation)
• Run preliminary assembly of extracted samples to assess DNA/RNA integrity, purity, genome size, heterozygosity and the presence of co-biont or contaminant material.
• Feed results back rapidly to the culturing, extraction and library preparation teams at EI and SAMS so that protocols can be iterated, and contribute the evidence underpinning the project’s lineage-aware extraction and lysis recommendations. - 15
- Contribution to the Institute
• Contribute to the wider scientific life of the Institute, including seminars, group meetings and cross-platform method development.• As agreed with line manager, any other duties commensurate with the nature of the role.
- 10
- Computational workflow development, packaging and reproducibility
• Ensure that all computational workflows developed on the project are made publically available• Support cross-laboratory benchmarking of workflows between EI, SAMS and the Bigelow SCGC, and troubleshoot deployment on partner infrastructure.
• Package analyses reproducibly (e.g. Nextflow/Snakemake, containers, version-controlled environments) so that methods are portable across the partner laboratories and readily adoptable by the wider community. - 10
- Data management, deposition and protocol dissemination
• Submit raw data, assemblies, annotations and transcriptome datasets to ENA and/or NCBI, annotated with rich metadata and made findable, accessible, interoperable and reusable (FAIR).• Maintain accurate records of analyses and outcomes to support project reporting and milestone tracking.
• Contribute computational protocols and documentation to the project’s protocols.io workspace and project website - 5
- Dissemination, community engagement and training
• Prepare results for publication in peer-reviewed journals and present the work at national and international meetings and & lead in the preparation of papers where appropriate.
• Contribute to the project’s community engagement programme, including the international launch workshop, the benchmarking hackathon and dedicated training events, and provide hands-on support to external participants applying project workflows to their own taxa.
• Contribute technical blog posts and updates to the project website and respond to community feedback on protocols and workflows. - 5
- As agreed with line manager, any other duties commensurate with the nature of the role
Person Profile
Education & Qualifications
- Requirement
- Importance
- PhD (awarded, or submitted/close to submission) in bioinformatics, computational biology, genomics, evolutionary biology or a closely related discipline
- Essential
- Undergraduate degree in a biological, computational, mathematical or other numerate discipline
- Essential
- Formal training or demonstrable self-directed expertise in eukaryotic genome assembly and annotation
- Desirable
- Training in software engineering, reproducible research or research data management
- Desirable
Specialist Knowledge & Skills
- Requirement
- Importance
- Practical, hands-on experience of de novo genome assembly using long-read data (PacBio HiFi and/or Oxford Nanopore)
- Essential
- Proficiency in at least one scripting/programming language used in bioinformatics (e.g. Python, R, Perl, C/C++, Rust)
- Essential
- Confident working in a Linux/UNIX command-line environment and running analyses on an HPC cluster (e.g. Slurm, LSF)
- Essential
- Working knowledge of assembly quality assessment, including BUSCO, k-mer based completeness and contiguity metrics
- Essential
- Experience of structural and/or functional annotation of eukaryotic genomes or transcriptomes
- Essential
- Competent use of version control (Git/GitHub) and a demonstrable commitment to reproducible, well-documented analysis
- Essential
- Ability to critically evaluate and benchmark competing computational methods and to communicate the results objectively
- Essential
- Experience with workflow management systems (e.g. Nextflow, Snakemake, Galaxy) and workflow sharing platforms such as WorkflowHub
- Desirable
- Experience of metagenomic or metagenome-adjacent analysis: read/assembly binning, decontamination and separation of co-biont genomes
- Desirable
- Experience of single-cell genomics and of the artefacts associated with whole-genome amplification (chimeras, uneven coverage, allelic dropout)
- Desirable
- Experience of scaffolding using long-range data (Hi-C, ultra-long ONT, linked or mapped-read technologies)
- Desirable
- Experience of long-read transcriptome analysis, isoform discovery and non-canonical gene structures
- Desirable
- Knowledge of protist / microbial eukaryote biology, phylogeny and genome diversity
- Desirable
- Familiarity with containerisation and environment management (Docker, Singularity/Apptainer, Conda)
- Desirable
Relevant Experience
- Requirement
- Importance
- Track record of analysing large-scale next-generation sequencing datasets from design through to interpretation
- Essential
- Publication record in peer-reviewed journals commensurate with career stage
- Essential
- Experience of working with genomic data from non-model organisms, or with genomes lacking high-quality references
- Essential
- Experience of troubleshooting analyses where the underlying data are imperfect, and of feeding conclusions back to laboratory colleagues to improve upstream protocols
- Essential
- Experience of submitting sequence data, assemblies and metadata to public archives (ENA, NCBI/GenBank)
- Desirable
- Experience of releasing open-source software, pipelines or public protocols used by others
- Desirable
- Experience of working as part of a multi-institution or international research consortium
- Desirable
- Experience of working alongside a sequencing or genomics service platform
- Desirable
- Experience of delivering computational training, workshops or hackathons
- Desirable
Interpersonal & Communication Skills
- Requirement
- Importance
- Demonstrated ability to work independently, using initiative and applying problem solving skills
- Essential
- Good interpersonal skills, with the ability to work as part of a multidisciplinary team spanning computational and laboratory scientists
- Essential
- Excellent communication skills, both written and oral, including the ability to present complex technical information with clarity to non-specialist audiences
- Essential
- Able to plan and prioritise a varied computational workload across three concurrent project aims and to meet agreed deadlines
- Essential
- Able to act as an ambassador for the Institute and build key relationships locally on the park and externally
- Essential
- Demonstrates a commitment to high personal and professional standards. Assumes responsibility and accountability for the successful completion of projects, assignments or tasks
- Essential
- Willingness and ability to support, mentor and train others in computational methods
- Desirable
Additional Requirements
- Requirement
- Importance
- Attention to detail
- Essential
- Willingness to embrace the expected values and behaviours of all staff at the Institute, ensuring it is a great place to work
- Essential
- Commitment to open science, FAIR data principles and the open release of code, workflows and protocols
- Essential
- Ability to maintain confidentiality and security of information where appropriate, including in relation to pre-publication data
- Essential
- Ability to undertake occasional national and international travel for project meetings, workshops, hackathons and conferences
- Essential
- Willingness to work outside standard working hours when required, including to accommodate meetings across international time zones
- Essential
- Promotes equality and values diversity
- Essential
- Able to present a positive image of self and the Institute, promoting both the international reputation and public engagement of the Institute
- Essential
Who We Are
Earlham Institute
About the Earlham Institute
The Earlham Institute harnesses the power of data-driven biology to accelerate solutions for health, biodiversity, and food security. Based at Norwich Research Park, the Earlham Institute is one of eight institutes strategically funded by BBSRC.
Our science combines world-class technology, interdisciplinary expertise, and training and development across genomics, engineering biology and data science, to decode the scale and complexity of living systems.
We believe we can achieve more if we work together. That's why we collaborate with the global science community and industry partners, while also inspiring the next generation of scientists and technical specialists.
Our Science
Earlham Institute scientists specialise in developing and testing the latest tools and approaches needed to decode living systems and make biological predictions.
We are home to state-of-the-art facilities and technology, creating a unique combination of expertise and infrastructure.
We have dedicated laboratories for genome sequencing, single-cell analysis, engineering biology, and large-scale automation; as well as one of the largest supercomputing facilities for life science research in Europe.
Our Advanced Training team also provides access to specialised scientific training to upskill the next generation of research and technical staff.
Our Culture
Our collegiate and innovative research environment comes with significant support, including a commitment to your professional development, research and administrative assistance, and opportunities to build collaborations with scientists and industry on the Norwich Research Park, across the UK, and internationally.
We are committed to building and maintaining a workplace that treats every individual with dignity and respect. By taking an active approach to fostering inclusivity, diversity, equality and accessibility, we empower our community to achieve more.
The Institute is also home to talented technical and operational staff, whose invaluable contributions enable our science to have the maximum impact. We aim to recognise, reward, and develop all staff and students so that every individual feels able to achieve their best with us.
We work hard to nurture an engaged and positive workplace, centred on core values that include openness, technical excellence, and collaboration. We attract staff from around the world who contribute to - and benefit from - an environment that enables them to deliver world-class science alongside a supportive and social community.
For more information about working at the Earlham Institute, please click here.
Further Information:
Department
Research Faculty
Group Details
The Earlham Institute is a BBSRC-supported research institute on the Norwich Research Park specialising in genomics, single-cell and spatial biology, and computational biology. The post sits within the Director's group and works across two of the Institute's core science platforms.
The Technical Genomics Group (Dr Karim Gharbi) operates the High-Throughput Sequencing platform and delivers genome and transcriptome data production, together with technical development for DNA/RNA isolation, library preparation and short- and long-read sequencing, including the evaluation of new and emerging methodologies.
The Single-Cell and Spatial Analysis platform (Dr Iain Macaulay) is one of the most advanced facilities in the UK for single-cell and spatial genomics of model and non-model organisms, and led environmental protist sequencing for the Darwin Tree of Life project. It develops imaging and spectral cell sorting approaches and long-read single-cell RNA sequencing.
The project team at EI additionally includes a dedicated project manager, experienced protist genomics and single-cell postdoctoral scientists, and senior research assistants covering DNA/RNA isolation and long-read library preparation. Externally, the consortium brings together CCAP/SAMS, who maintain one of the world's largest and most taxonomically diverse protist culture collections (~3,200 strains), and the Bigelow Laboratory Single Cell Genomics Center, the world's first single-cell genomics centre focused on environmental microorganisms.
Living in Norfolk
Advertisement
Postdoctoral Scientist (Bioinformatics) - Protist Genomics
The Earlham Institute has been awarded funding by the Gordon and Betty Moore Foundation to develop Protist Omics at Scale, a three-year international methods-development programme run in partnership with the Scottish Association for Marine Science, home of the Culture Collection of Algae and Protozoa, and Aalborg University.
We are looking for a computational biologist to lead the computational core of the project: quality control, assembly, decontamination and co-biont separation, and structural and functional annotation across all three aims. Where existing tools fail, the postholder will diagnose why and develop what replaces them.
The post is based at the Earlham Institute on the Norwich Research Park.
Background:
Protists represent the vast majority of eukaryotic diversity but remain significantly under-represented in reference genome databases. Their genomes are often large, repetitive and genetically complex, and are frequently derived from mixed, low-biomass or uncultured samples, making them difficult to assemble and annotate using standard genomic approaches.
This project aims to address these challenges by systematically identifying and overcoming key bottlenecks in genome and transcriptome assembly from bulk cultures and single cells.
Based within the Earlham Institute's Director's Group, the project combines expertise in long-read sequencing, single-cell genomics, spatial biology and computational biology. It brings together leading facilities at the Earlham Institute, including the Technical Genomics Group and the Single-Cell and Spatial Analysis Platform, as well as external collaborators at CCAP/SAMS, home to one of the world's largest protist culture collections, and Aalborg University.
The overall objective is to develop and apply innovative methods that enable the generation of high-quality genomic and transcriptomic resources for previously inaccessible and poorly characterised eukaryotic organisms.
The role:
This is a postdoctoral computational biology/bioinformatics role focused on developing and applying novel methods for long-read and single-cell genome and transcriptome assembly across a diverse range of protist species.
The postholder will:
• Develop expertise in advanced genome and transcriptome assembly approaches.
• Work on complex long-read and single-cell sequencing datasets.
• Contribute to the development of new computational methods rather than routine analysis.
• Develop research software engineering skills, including workflow development, packaging and containerisation.
• Create and maintain reproducible bioinformatics workflows using platforms such as Galaxy and WorkflowHub.
• Collaborate closely with internal and external partners across the consortium.
• Lead or contribute significantly to project outputs.
• Publish research findings and present at national and international conferences.
• Participate in workshops, hackathons and community training activities.
• Support the supervision and development of students where appropriate.
The role offers extensive opportunities for career development, networking and collaboration within an internationally recognised genomics research environment.
The ideal candidate:
The post holder will have, or be close to completing, a PhD in bioinformatics, computational biology, genomics, evolutionary biology or a closely related discipline.
They will have practical experience of analysing large-scale next-generation sequencing datasets and de novo genome assembly using long-read sequencing data (PacBio HiFi and/or Oxford Nanopore), together with proficiency in at least one bioinformatics programming language and experience working in a Linux/HPC environment.
The successful candidate will have experience of genome or transcriptome analysis, an ability to critically evaluate computational methods, and a track record of contributing to research outputs, including peer-reviewed publications.
Experience of workflow development and reproducible research practices, including version control and workflow management systems, would be advantageous, as would knowledge of single-cell genomics, protist or microbial eukaryote biology, and software containerisation technologies.
Additional information:
This is a full-time post for a contract of 36 months.
Salary on appointment will be within the range £39,000 - £46,500 per annum, depending on qualifications and experience. A starting salary of £40,100 is guaranteed for candidates who can evidence their PhD certificate at appointment; those awaiting confirmation of their PhD award will be appointed at £39,000 until evidence is provided.
This role meets the criteria for a visa application, and we encourage all qualified candidates to apply. Please contact the Human Resources Team if you have any questions regarding your application or visa options.
As a Disability Confident employer, we guarantee to offer an interview to all disabled applicants who meet the essential criteria for this vacancy.
The closing date for applications will be 1 October 2026.