Models & research
Hugging Face
Introducing ConTextual: How well can your Multimodal model jointly reason over text and image in text-rich scenes?
This source did not provide an excerpt. Read the announcement at the publisher.
Excerpt supplied by the publisher’s feed