PhD Researcher - ENSI Manouba - 2023 - Present
A sequence-to-sequence protein language model bridging sequence understanding and generation, currently in resubmission to ISMB/ECCB, a Bioinformatics journal, and NeurIPS GenBio.
Structured UniProt into a queryable biomedical knowledge graph to power instruction-dataset construction for LLM fine-tuning.
A controlled tokenizer ablation on a 120M-parameter causal language model trained on SwissProt sequences, comparing amino-acid-level vs. 10K/20K/32K BPE vocabularies to isolate the effect of domain-specific vocabulary design on downstream protein tasks.
Designing a chain-of-thought pretraining signal grounded in static biophysical features, plus a reinforcement-learning stage with structural/functional reward models.
Distilled a 1.3B-parameter teacher (InstructProtein) into a 400M-parameter student via alternating layer extraction (6L-ALT), then fine-tuned on a 212K-example in-house instruction dataset built from SwissProt.
Ran GRPO (RLVR) on the SFT checkpoint with a composite verifiable reward - format gate, PSIPRED-based secondary-structure correctness, and a Gaussian length penalty - on Colab Pro A100.