ID mapping based on seq_block.xlsx
- level column: high level to low level
- seq_len: longer sequence to shorter sequence
- print out unmatched sequence if any, get their info and assign their ID
for 'duplicate' seq, copy mapped ID to fill the blank
type mapping from ID_* columns
replace ID with type in seq_block.xlsx
backbone mapping from type_* columns
- merge from left right (important)
- follow the merging rules:
VH, CH1 = FabH VL, CL = FabL
VH, linker, VL = scFv_HL VL, linker, VH = scFv_LH
hinge, CH2CH3 = Fc
content mapping from backbone_* columns
- follow merging rules scFv_HL = scFv scFv_LH = scFv FabH + FabL = Fab
Fc x2 = Fc, if Fc is odd number, put Fc? (PS: if seeing Fc?, meaning chain number is wrong)
protein example should be |<protein name> prot| (e.g., NRP1_b1b2 prot)
- infer target/protein from ID_* columns, put in front of backbone type
- always put Fc behind others
- if target/protein + type occur twice, then use xNumber format,e.g., EGFR scFv x2