initial tuned 4k context length model

This is a continued pretrained version of Florence-2-large model with 4k context length with original data, only 0.1B samples are used for continue pretraining, thus it might not be trained well. In addition, OCR task has been updated with line separator ('\n'). This model has COCO OD AP 39.8. The model is a starting point for experimenting longer context length (4k) and is not aimed to replace the original model.

haipingwu

Microsoft org Nov 16, 2024

load the model with the following code

device = "cuda:0" if torch.cuda.is_available() else "cpu"
torch_dtype = torch.float16 if torch.cuda.is_available() else torch.float32
model = AutoModelForCausalLM.from_pretrained('microsoft/Florence-2-large', torch_dtype=torch_dtype, trust_remote_code=True, revision='refs/pr/84').to(device)
processor = AutoProcessor.from_pretrained('microsoft/Florence-2-large', trust_remote_code=True, revision='refs/pr/84')

haipingwu pinned discussion Nov 16, 2024

anhkr

Nov 20, 2024

@haipingwu
Could you tell us the steps you took to extend the token length from 1024 to 4096? I would like to understand the process so I can apply it to extend the token length for Florence-2-base-ft in my task.

haipingwu

Microsoft org Nov 20, 2024

hi @Ank12 , you can check the change of config.json, where max_position_embeddings has been changed from 1024 to 4096. Then you can fine-tune the model with the new configuration. For the extended new embedding weight init, you can either init from scratch or interpolate it with the pre-trained weights.

leoxiaobin changed pull request status to open Dec 8, 2024

leoxiaobin changed pull request status to merged Dec 8, 2024

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment