Genomics & Bioinformatics Server Guide for Indian Labs
By Rohit, Founder ·

Quick answer: A genomics or bioinformatics compute server needs many CPU cores for alignment and variant calling, large ECC memory (the BWA-MEM2 aligner alone needs roughly 28 × the reference size in RAM to build a human genome index, about 84 GB), and a lot of fast storage, because a single 30x human genome typically takes 100–200 GB as raw FASTQ and 80–100 GB as an aligned BAM. A GPU is optional: it only helps if you run GPU-accelerated tools such as NVIDIA Parabricks, which needs at least 16 GB of GPU memory and, for a 2-GPU system, at least 100 GB of system RAM and 24 CPU threads.
This guide is for university labs, research institutes, diagnostic and biotech teams in India who are deciding what to buy to run sequencing pipelines in-house. It is deliberately about the pipeline itself — which stage needs what — rather than a generic "research server" list. For the wider question of equipping a shared college lab, see our college research lab compute server guide; for deep-learning workloads, see the AI / ML server page.
What a Typical Sequencing Pipeline Asks of Hardware
Most short-read germline pipelines follow the same broad stages. Each stage stresses a different part of the machine, which is why one number ("how many cores?") never answers the buying question:
| Pipeline stage | Example tools | Main hardware pressure |
|---|---|---|
| Quality control and trimming | FastQC, fastp | Storage read speed, a few cores |
| Alignment to a reference genome | BWA-MEM2, minimap2 | CPU threads; RAM (especially for building indexes) |
| Sorting, duplicate marking, indexing | samtools, GATK, Picard | Disk throughput and temporary space; RAM |
| Variant calling | GATK HaplotypeCaller, DeepVariant | Many independent jobs run in parallel; RAM per job |
| Joint genotyping and annotation (cohorts) | GATK, annotation databases | RAM and fast storage for many samples at once |
The tool names above are widely used examples, not a recommendation of a specific workflow — your lab's protocol, reference genome and sequencing platform decide the real pipeline. The hardware logic holds either way.
CPU: Core Count Matters Because the Work Splits Up
Alignment tools are built to use many threads. BWA-MEM2's documentation shows multi-threaded runs with a -t thread-count flag, and its published benchmarks were run on a 56-core system. Variant calling scales differently. Broad Institute's documentation describes scatter-gather parallelism: the computation is split into smaller independent pieces (for example genomic intervals, or separate samples), those pieces run in parallel, and the results are gathered into one. That design is why a machine with a high core count and enough memory per job can process many samples at once, and why a few very fast cores are not a good substitute.
In practice this points to a many-core server CPU. ProStation builds on Intel Xeon Scalable and AMD EPYC platforms; our EPYC vs Xeon guide explains how to choose between them, and the CPU vs GPU balance guide helps decide how much of your budget belongs on each side. If your throughput target is many genomes per week rather than one at a time, a dual-CPU platform is the usual direction — our Ultra tier is built on dual Intel Xeon Scalable or AMD EPYC.
RAM: Index Building and Parallel Jobs Set the Floor
Memory is where genomics most often surprises first-time buyers. Two documented examples:
- Building a BWA-MEM2 index: the project's documentation states it requires 28N GB of memory, where N is the size of the reference sequence in gigabases. For the human genome (about 3 Gb) that is roughly 84 GB. You do this once per reference, but the machine must be able to do it.
- GPU-accelerated pipelines: NVIDIA's Parabricks documentation lists minimum system RAM and CPU threads by GPU count — at least 100 GB RAM and 24 threads for 2 GPUs, 196 GB and 32 threads for 4 GPUs, and 392 GB and 48 threads for 8 GPUs.
On top of those floors, every variant-calling job running in parallel needs its own working memory, so total RAM scales with how many jobs you want in flight. Genomics runs are long, and a silent memory error in hour 20 of a pipeline is expensive; this is the situation ECC memory is designed for — see ECC vs non-ECC RAM. All ProStation Starter, Pro and Ultra tiers use ECC memory.
Storage: Plan for Accumulation, Not Just One Sample
Sequencing data is large and it piles up. The planning ranges below are representative figures published by Lab Manager for a 30x human whole genome; actual sizes vary with platform, read length, coverage and pipeline.
| File type | Approx. size per 30x human genome | Role |
|---|---|---|
| FASTQ (raw reads) | ~100–200 GB | Starting input |
| BAM (aligned) | ~80–100 GB | Working copy for analysis and re-analysis |
| CRAM (compressed alignment) | ~30–60 GB | Long-term archive; lossless for the aligned data |
| VCF (variants) | ~1 GB | Final calls |
Lab Manager cites roughly 250 GB per genome as a common full-retention planning figure and notes that CRAM cuts aligned-file size by about 40–70%. A lab sequencing 20 genomes a month is therefore planning for terabytes within months, not years. A sensible layout has three tiers:
- Fast scratch (NVMe SSD): for the alignment, sorting and calling stages, where intermediate files are large and disk speed shows up directly in run time. See NVMe vs SAS vs SATA.
- Bulk working storage (HDD array): for projects in progress, protected by RAID — see RAID levels explained. Remember that RAID protects against a drive failure, not against deletion or corruption.
- Archive and backup: a separate copy, ideally on a different machine — a storage and backup server or a NAS as described in our NAS setup guide. Converting finished BAMs to CRAM is the main lever for keeping this tier affordable.
Do You Need a GPU?
Not by default. The mainstream CPU-based tools above need none. A GPU becomes useful if your team adopts GPU-accelerated pipelines. NVIDIA Parabricks, for example, requires at least 16 GB of GPU memory per GPU, with 18 GB or more needed for its default HaplotypeCaller configuration, supports specific CUDA architectures (listed in NVIDIA's installation requirements — check your GPU against it), and states that a single GPU is supported but not recommended. Because the requirements depend on the exact tool and version, confirm them against NVIDIA's current documentation before choosing a card. If you also run machine-learning work on the same machine, the GPU decisions in our shared GPU workstation guide apply.
Networking, Power and Where It Lives
- Network: sequencers and analysis servers move very large files. If data lands from an instrument or a shared store, 10GbE is worth considering over 1GbE; the networking guide shows the bandwidth math.
- Power: a pipeline that runs for many hours should not die in a power cut. See the UPS sizing guide.
- Form factor: a single lab usually fits a tower server in a room; larger groups may prefer a rack — compare them in the tower vs rack guide.
- On-premise vs cloud: if sequencing volume is steady, owning the hardware is often easier to plan; irregular, bursty projects may suit cloud. Our where should your server live guide covers the trade-off. Human genomic data is sensitive, so also check your institution's ethics and data-governance rules about where such data may be stored.
Starting Points by Lab Size
ProStation does not publish pipeline benchmarks, and we do not claim specific genomes-per-day figures. What we can do is map the requirements above onto our existing build tiers, which are listed on the servers page:
| Lab profile | Reasonable tier to start discussing | Why |
|---|---|---|
| Small group, a few samples at a time, exploratory work | Starter (Xeon E / EPYC entry, 16–64 GB ECC, 1 TB NVMe) | Fine for learning and small targeted panels; 84 GB index builds and whole-genome cohorts will need more RAM and storage than the base spec, so ask for a custom configuration |
| Active research group, regular whole-genome or exome batches | Pro (Xeon Scalable / EPYC Milan, 64–256 GB ECC, 2 TB NVMe + HDD options) | RAM clears the index-build floor and supports several parallel jobs; HDD options for bulk working data |
| Core facility, cohort studies, GPU-accelerated pipelines | Ultra (dual Xeon Scalable / EPYC Genoa, 256–512 GB ECC, 4 TB+ NVMe RAID, NVIDIA RTX / A-series / Tesla GPU) | Dual-CPU thread count and RAM headroom for multi-GPU system-RAM requirements |
These are starting points for a conversation, not final specifications — the right build depends on your pipeline, sample volume and retention policy.
Why Choose ProStation Systems for a Research Compute Server
- Built to your pipeline, not a fixed SKU: every build is custom, so RAM, scratch NVMe, bulk HDD and GPU can be balanced to your stages. Start with the configurator or free pre-buy consulting and share your tools, reference genome and monthly sample count.
- Brand-new hardware with a warranty: long pipelines make downtime costly. Our warranty tiers and the downtime cost guide show how to weigh that. Builds are delivered in about four working days.
- Education and research focus: our Education & Research page lists bioinformatics among the workloads we build for. Teams in Hyderabad's life-sciences cluster can see local details on our Hyderabad page.
- Honest advice on budgets: if a department only needs a lower-cost secondary or archive machine, new vs refurbished explains the trade-offs, and our sister brand Serverwale sells tested refurbished servers. For the primary analysis machine that runs for days at a time, most labs prefer new hardware under warranty.
"We needed servers for our hospital management system. ProStation Systems handled everything — consultation, delivery, installation, and even trained our IT staff. Very professional and reliable."
— Sunita Patel, Admin, Government Medical College, Gujarat
That project was a hospital system rather than a sequencing lab, but the process — consultation first, then delivery and installation — is the same one we follow for research compute.
Checklist Before You Buy
- List the tools and versions in your pipeline, and your reference genome.
- Estimate samples per month and the average size of each (FASTQ, BAM/CRAM, VCF).
- Decide your retention policy: how long BAMs are kept, and when they are converted to CRAM.
- Size RAM from the largest single step (for example index building) and from the number of parallel jobs.
- Decide whether you will run GPU-accelerated tools, and check the exact GPU memory and architecture requirements first.
- Plan three storage tiers: fast scratch, bulk working storage and a separate backup.
- Check network speed to your data source, power backup and your institution's data-governance rules.
Frequently Asked Questions
Q1. How much RAM does a genomics server need?
It depends on the tool. BWA-MEM2's documentation states that building a human genome index needs about 28 × the reference size in GB, roughly 84 GB. NVIDIA Parabricks lists at least 100 GB of system RAM for a 2-GPU system. Parallel variant-calling jobs add working memory on top, so size RAM for the largest step plus the number of jobs you want to run at once.
Q2. How much storage does one whole genome need?
Representative planning ranges for a 30x human genome are about 100–200 GB as FASTQ, 80–100 GB as BAM, 30–60 GB as CRAM and about 1 GB as VCF. Full retention is often planned at around 250 GB per genome. Actual sizes vary with platform and pipeline.
Q3. Do I need a GPU for bioinformatics?
No. Common CPU-based tools such as BWA-MEM2 and GATK do not need one. A GPU helps only if you adopt GPU-accelerated pipelines such as NVIDIA Parabricks, which requires at least 16 GB of GPU memory per GPU and recommends more than one GPU.
Q4. Is ECC memory necessary for genomics?
It is not a software requirement, but long multi-hour runs on large datasets are the case ECC is designed for, since it detects and corrects single-bit memory errors. All ProStation server tiers use ECC RAM.
Q5. Should we buy a server or use cloud for sequencing analysis?
Steady, predictable sample volume often suits owned hardware; irregular bursts may suit cloud. Because human genomic data is sensitive, also follow your institution's rules on where it can be stored. Our on-premise vs colocation vs cloud guide covers the trade-offs.
Q6. Can ProStation build a custom server for our bioinformatics lab?
Yes. ProStation Systems builds brand-new custom servers and workstations, and bioinformatics is one of the workloads listed on our Education & Research page. Share your pipeline, sample volume and budget through free consulting or the configurator and we will recommend a configuration.
Final Recommendation
Size a genomics server from the pipeline outward: many-core CPUs for alignment and parallel variant calling, ECC RAM above your largest single step, three tiers of storage with a real backup, and a GPU only if your tools use one. Write down your sample volume and retention policy before you choose parts — storage growth, not CPU speed, is what most labs underestimate. To have a configuration specified for your lab, call ProStation Systems on +91 87968 22044, message us on WhatsApp, or contact us.
Sources: BWA-MEM2 documentation; NVIDIA Parabricks installation requirements; Lab Manager — sizing storage for NGS data; Terra / Broad Institute — scatter-gather parallelism.