Skip to content

Questions on creating instruction data #13

Open
@henryhungle

Description

@henryhungle

Thanks for the great work!

I have a few questions regarding data creation of xP3 after following the guide here to create instruction data on the code language subset.

  1. I noticed the total samples of the public processed data (from here) on the code split is 2707724. However, my resulting data following the above github guide is much more than that (approximately >3M samples). I wonder if there were any additional post-processing to get the final instruction data for tuning?

  2. Following the above github guide, I noticed there was no prompt for this particular dataset State Changes. I got this warning when running the creation code:
    Tried instantiating `DatasetTemplates` for Fraser/python-state-changes, but no prompts found. Please ignore this warning if you are creating new prompts for this dataset.

Is this dataset not assigned with any prompt (similar to how HumanEval was treated). Or is the below version of PromptSource I used is not correct:
git clone -b tr13 https://github.com/Muennighoff/promptsource.git & install cd promptsource; pip install -e .

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions