> For the complete documentation index, see [llms.txt](https://ubiai.gitbook.io/ubiai-documentation/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://ubiai.gitbook.io/ubiai-documentation/upload-documents.md).

# Upload Documents

In order to **upload** a new document, click on project list in the project drop down menu and click on cloud upload button backup. Different formats are accepted:

1. TXT,PDF, HTML and DOCX
2. Native PDF
3. JPG, PNG
4. AWS OCR Textract (ZIP file containing JSON textract output and its associated PDF or image file)
5. JSON: you can upload a JSON file with existing entities and relations. This is useful if you have a pre-annotated JSON file and you don’t want to re-annotate.\
   The JSON should follow the format below:

<figure><img src="https://3073024999-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F0hqV0hLtifjaqlfkGT7n%2Fuploads%2F38OfTi4mbTEy7CPgh3Ko%2Fjson_format.png?alt=media&amp;token=38170d33-910f-4e7d-8983-7eb483dc1f53" alt=""><figcaption></figcaption></figure>

6. CSV (UTF-8): you can upload a csv file containing one document per row. Note that the CSV needs to be UTF-8 encoded.

<figure><img src="https://3073024999-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F0hqV0hLtifjaqlfkGT7n%2Fuploads%2FnQnpwZUa7zEGvgx7dH9U%2Fcsv_list.png?alt=media&amp;token=8423b4c8-0337-422a-a8e6-9710079b474a" alt=""><figcaption></figcaption></figure>

7. TSV: you can upload a .tsv file containing pre-tagged tokens following the IOB format as shown below. Documents are seperated by the token -DOCSTART- -X- O O at the start of each document.

<figure><img src="https://3073024999-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F0hqV0hLtifjaqlfkGT7n%2Fuploads%2FcUs3KV6Wlkv5mxtsqszC%2Ftsv_upload%20(1).png?alt=media&amp;token=9969fc09-c223-48b0-a06d-6b32525bd0bb" alt=""><figcaption></figcaption></figure>

8. ZIP: you can upload a zip file containing TXT, PDF or HTML. This is useful to upload documents in bulk.

<figure><img src="https://3073024999-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F0hqV0hLtifjaqlfkGT7n%2Fuploads%2Fm619DIW4PFAabdEuaYMO%2Fupload.png?alt=media&amp;token=a5d9ecc4-89ff-40f8-823b-b122ca19aea8" alt=""><figcaption></figcaption></figure>

Once you select your documents, you will then be prompted to choose between 3 pre-annotation options:

* Dictionary: Auto-annotate the uploaded document(s) using the project's dictionary (not applicable for JSON and TSV upload)
* Models: Auto-annotate the uploaded document(s) using your trained ML model
* No pre-annotation: No auto-annotation

You also have the option to remove duplicate documents during upload by checking "remove duplicate documents".

<figure><img src="https://3073024999-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F0hqV0hLtifjaqlfkGT7n%2Fuploads%2FjzGQxJXj0NZCRZJyCbi1%2Fimg_pre_ann2.png?alt=media&amp;token=b5d4c164-a055-4c9b-be35-7eca239a68a5" alt=""><figcaption></figcaption></figure>

For native PDF, JPG and PNG using OCR, you have the option to choose between three type of OCR engine:

* Default: Engine will be chosen based on the language of the project
* OCR 1: Engine based on AWS Textract
* OCR 2: Engine based on Google Vision API
* OCR 3: Engine based on Microsoft Azure OCR, includes the option to automatically parse tables

<figure><img src="https://3073024999-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F0hqV0hLtifjaqlfkGT7n%2Fuploads%2Fx67xsigUSFrZuQYKqHsb%2Fupload-ocr.PNG?alt=media&amp;token=7f919acf-c72c-4cb4-a6a7-0f320aa7e57f" alt=""><figcaption></figcaption></figure>
