My Research

Translating the Language of Life

PhD Researcher - ENSI Manouba - 2023 - Present

Life2Lang

A sequence-to-sequence protein language model bridging sequence understanding and generation, currently in resubmission to ISMB/ECCB, a Bioinformatics journal, and NeurIPS GenBio.

UniprotKG

Structured UniProt into a queryable biomedical knowledge graph to power instruction-dataset construction for LLM fine-tuning.

Eshmun

A controlled tokenizer ablation on a 120M-parameter causal language model trained on SwissProt sequences, comparing amino-acid-level vs. 10K/20K/32K BPE vocabularies to isolate the effect of domain-specific vocabulary design on downstream protein tasks.

Designing a chain-of-thought pretraining signal grounded in static biophysical features, plus a reinforcement-learning stage with structural/functional reward models.

Eshmun-R1

Distilled a 1.3B-parameter teacher (InstructProtein) into a 400M-parameter student via alternating layer extraction (6L-ALT), then fine-tuned on a 212K-example in-house instruction dataset built from SwissProt.

Ran GRPO (RLVR) on the SFT checkpoint with a composite verifiable reward - format gate, PSIPRED-based secondary-structure correctness, and a Gaussian length penalty - on Colab Pro A100.

Research Objectives

  • Develop protein language models that unify sequence understanding and generation
  • Build structured knowledge resources (UniprotKG) to support LLM fine-tuning on biological data
  • Study how domain-specific tokenization and vocabulary design affect downstream protein tasks
  • Apply reasoning-oriented post-training (chain-of-thought, RLVR) grounded in biophysical and structural reward signals

Publications