User Manual

Complete guide to using the ShEnrich - Shrimp Enrichment Analysis Platform

ShEnrich User Guide

Everything you need to know about this platform

Getting Started

Welcome to the ShEnrich - Shrimp Enrichment Analysis Platform. This comprehensive resource provides curated metabolic pathway information, protein sequences, and functional annotations for five major shrimp species used in aquaculture research.

847 Metabolic Pathways
95,504 Protein Sequences
8,634 Nutrient Compounds
5 Shrimp Species

Supported Species

Scientific Name Common Name Code Proteins
Penaeus vannamei Pacific white shrimp Pvan 25,154
Penaeus monodon Giant tiger prawn Pmon 23,892
Penaeus chinensis Chinese shrimp Pchi 19,876
Penaeus indicus Indian white prawn Pind 14,323
Penaeus japonicus Japanese tiger prawn Pjap 12,259
Quick Start Guide
  1. Search by ID: Navigate to Explore → Search by ID for quick protein information lookup using XP IDs
  2. BLAST Search: Use Explore → Search by Sequence to find similar proteins across species
  3. Batch Annotation: Access Explore → Annotate to retrieve comprehensive functional annotations for multiple proteins
  4. Pathway Enrichment: Use Enrich → Pathway Enrichment to identify overrepresented KEGG pathways in your gene lists
  5. GO Enrichment: Navigate to Enrich → GO Enrichment for Gene Ontology term enrichment analysis
  6. Genome Visualization: Explore Explore → Genome Browser to visualize genomic features interactively

Explore Tools

The Explore menu provides comprehensive tools for searching, browsing, and analyzing shrimp genomic data. Access these tools from the top navigation: Explore → [Tool Name]

Search by ID
Quickly retrieve protein information using accession IDs. Supports XP_ format protein IDs with or without version numbers.
Example: XP_047494425.1, XP_027225530.1
Access: Explore → Search by ID
Search by Sequence
BLAST search tool to find similar proteins across all species. Submit protein sequences in FASTA format and get results with E-values and alignment scores.
Input: FASTA format
>Query_protein
MTEITAAMVKELRESTGAGMMDCK...
Access: Explore → Search by Sequence
Annotate
Batch annotation lookup for multiple proteins. Retrieve GO terms, InterPro domains, enzyme codes, KEGG pathways, and protein descriptions.
Input: List of XP IDs (one per line)
Supports file upload or paste
Access: Explore → Annotate
Find Orthologs
Search for orthologous genes across shrimp species or browse gene family groups. Essential for comparative genomics and evolutionary analysis.
Search by: Protein ID, Gene Family ID, or Species filter
Access: Explore → Find Orthologs
Genome Browser
Interactive JBrowse-powered visualization of genomic features. View gene locations, pathway annotations, and chromosomal features with zoom and pan capabilities.
Features: Gene tracks, pathway tracks, search by region or gene
Access: Explore → Genome Browser

Search Tips & Best Practices

Optimization Tips
  • ID Format: Use XP_ format IDs (e.g., XP_047494425.1). Version numbers (.1, .2) are optional
  • BLAST E-value: Default 1e-5 works for most searches. Lower for more stringent matches
  • Batch Annotations: Can process up to 1000 protein IDs at once via Annotate tool
  • Species Codes: PC=P.chinensis, PI=P.indicus, PM=P.monodon, PJ=P.japonicus, PV=P.vannamei
  • Download Results: All search results can be exported in CSV, TSV, or JSON formats

Metabolism Categories Covered

Carbohydrate Metabolism
Glycolysis, gluconeogenesis, pentose phosphate pathway, starch and sucrose metabolism, citrate cycle (TCA)
Lipid Metabolism
Fatty acid synthesis, β-oxidation, sterol biosynthesis, phospholipid metabolism, sphingolipid metabolism
Amino Acid Metabolism
Essential amino acid biosynthesis, degradation pathways, amino acid derivatives, nitrogen metabolism
Cofactor & Vitamin Metabolism
Biosynthesis of enzyme cofactors, coenzyme metabolism, vitamin precursors, metal ion cofactor handling
Other Metabolism
Nucleotide metabolism, secondary metabolites, xenobiotics biodegradation, energy metabolism

Enrichment Tools

The Enrich menu provides statistical tools for functional enrichment analysis. Identify over-represented biological pathways and GO terms in your gene lists. Access these tools from the top navigation: Enrich → [Tool Name]

Pathway Enrichment
Statistical analysis to identify overrepresented KEGG metabolic pathways in your gene list. Uses hypergeometric test with multiple testing correction.
Input: Gene list (XP IDs, one per line)
Example:
XP_047494425.1
XP_027225530.1
XP_027238455.1
Access: Enrich → Pathway Enrichment
GO Enrichment
Analyze Gene Ontology term enrichment in three categories: Biological Process (BP), Cellular Component (CC), and Molecular Function (MF).
Input: Gene list (XP IDs, one per line)
Select: GO categories (BP/CC/MF)
Access: Enrich → GO Enrichment

Running Enrichment Analysis

Step-by-Step Process
  1. Prepare gene list: Collect protein accession numbers (XP_ format) from your differential expression or other analysis
  2. Access tool: Navigate to Enrich → Pathway Enrichment or Enrich → GO Enrichment
  3. Upload data: Paste gene IDs in text area OR upload a text file (one ID per line)
  4. Select species: Choose the appropriate shrimp species from dropdown menu
  5. Set parameters:
    • P-value cutoff (default: 0.05)
    • Multiple testing correction method (Benjamini-Hochberg recommended)
    • For GO: Select categories to analyze (BP/CC/MF)
  6. Run analysis: Click "Analyze Enrichment" and wait for results (typically 1-3 minutes)
  7. Interpret results: Review enriched pathways/terms with adjusted p-values < 0.05
  8. Visualize & Export: Use built-in plots and download options

Understanding Results

Result Columns Explained
Column Description
Pathway/Term ID KEGG pathway ID (e.g., ko00010) or GO term ID (e.g., GO:0006096)
Pathway/Term Name Human-readable name of the pathway or GO term
Gene Count Number of your input genes found in this pathway/term
Total in Pathway Total genes annotated to this pathway/term in the database
P-value Statistical significance (lower = more significant)
Adjusted P-value P-value after multiple testing correction (use this for interpretation)
Gene IDs List of your genes annotated to this pathway/term

Visualization Options

Bar Plot
Classic bar chart showing top enriched pathways ranked by significance. Bars colored by adjusted p-value.
Dot Plot
Dot size represents gene count, color represents p-value. Excellent for comparing multiple dimensions.
Network Plot
Shows relationships between enriched pathways and genes. Helps identify pathway crosstalk.
Export Options
Download results table (CSV/TSV) and publication-ready plots (PNG, PDF, SVG formats).

Best Practices for Enrichment Analysis

Important Considerations
  • Input size: Optimal gene list size is 50-500 genes. Too few (<20) may lack power, too many (>1000) may lack specificity
  • Background: All genes in the database are used as background for statistical testing
  • Multiple testing: Always use adjusted p-values (FDR) for interpretation, not raw p-values
  • Significance cutoff: Adjusted p-value < 0.05 is standard, but can be adjusted based on your study
  • Biological validation: Enrichment results should be validated with additional experiments or literature
  • Species-specific: Ensure your genes match the selected species database
Tips for Better Results
  • Quality control: Remove duplicate IDs and verify all IDs are valid before analysis
  • Fold-change filtering: Use biologically meaningful fold-change cutoffs (e.g., |log2FC| > 1) before enrichment
  • P-value selection: Start with adjusted p-value < 0.05, then explore borderline results (0.05-0.10)
  • Compare categories: Run both Pathway and GO enrichment for comprehensive functional insights
  • Pathway visualization: Use network plots to understand relationships between enriched pathways

Data Download

ShrimpEnrich provides multiple download options for bulk data access, analysis results, and reference datasets to support your research workflow.

Bulk Data Downloads
Access complete proteomes, pathway annotations, and GO terms for all species via Browse → Download Data section.
Search Results Export
Export search results and analysis outputs in CSV, Excel, or TSV formats for further computational analysis.
Analysis Results
Download enrichment analysis results, visualization plots, and gene lists in publication-ready formats.

Available File Formats

Data Type Format Best Use File Size
Protein Sequences FASTA Sequence analysis, BLAST databases 10-50 MB per species
Pathway Annotations GMT, TSV Enrichment analysis software 1-5 MB
GO Annotations GMT Gene Ontology analysis tools 2-8 MB
Search Results CSV, Excel Spreadsheet analysis Variable
Ortholog Groups TSV Comparative genomics 15-25 MB
Download Guidelines
  • Large files: Complete proteomes may take several minutes to download
  • File integrity: Check file sizes match expected values after download
  • Citing data: Please cite ShrimpEnrich when using downloaded data in publications
  • Updates: Data is updated quarterly - check version dates for currency

Troubleshooting

Search Problems

No search results found

  • Check spelling and try alternative protein names or synonyms
  • Use broader search terms (e.g., "kinase" instead of "protein kinase A")
  • Try searching without species filters to see if protein exists in other species
  • Remove special characters and use only letters, numbers, and spaces

Search results seem incomplete

  • Try "All Species" instead of specific species selection
  • Check if metabolism filter is too restrictive
  • Use Advanced Search for more comprehensive filtering options
BLAST Issues

BLAST search fails or hangs

  • Ensure sequence is in proper FASTA format with header line starting with >
  • Use only standard amino acid codes (A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y)
  • Remove any numbers or special characters from sequence
  • Try with a shorter sequence if original is very long (>2000 amino acids)

No BLAST hits returned

  • Lower E-value threshold or try more relaxed parameters
  • Check if query sequence is complete and not fragmentary
  • Verify sequence is from a eukaryotic organism (prokaryotic sequences may not match)
  • Try BLASTp with longer query sequences (minimum 30 amino acids recommended)
Enrich Analysis Problems

Analysis returns no significant results

  • Increase p-value threshold (try 0.1 instead of 0.05)
  • Ensure gene list contains at least 20-50 genes for robust analysis
  • Check that gene IDs match the selected species database
  • Try both pathway and GO enrichment - some gene sets may only be significant in one analysis

Gene IDs not recognized

  • Use protein accession numbers (XP_ format) rather than gene symbols
  • Ensure IDs are from supported species (P. vannamei, P. monodon, etc.)
  • Remove version numbers from accessions (e.g., use XP_047494425 instead of XP_047494425.1)
  • Check for extra spaces or characters in gene list
Download & Browser Issues

Downloads fail or are corrupted

  • Check browser download settings and available disk space
  • Disable popup blockers and ad blockers temporarily
  • Try downloading smaller data subsets if full datasets fail
  • Use a different browser (Chrome or Firefox recommended)

Pages load slowly or incompletely

  • Enable JavaScript in browser settings
  • Clear browser cache and cookies for the site
  • Disable browser extensions that might interfere
  • Check internet connection stability
Getting Additional Help

If problems persist after trying these solutions:

  1. Check system requirements: Modern browser with JavaScript enabled
  2. Document the issue: Note specific error messages and steps to reproduce
  3. Contact support: Email contact@ciba.res.in with details
  4. Include information: Browser type/version, operating system, screenshot if helpful
Response Times & Maintenance
  • Search queries: Usually complete within 5-10 seconds
  • BLAST searches: Typically 1-3 minutes depending on sequence length
  • Enrichment analysis: Generally 30 seconds to 2 minutes