About the problem formulation and evaluation results

Thanks for your great work!

I am wondering why GPT-Driver based on LLM outperforms the systems designed specifically for the motion planning task (e.g., NMP & UniAD, table I). If I understand correctly, it seems that you only provide limited information to LLM, while some important driving context is ignored (e.g., lane structure, traffic signals, etc.). Besides, when taking the visual grounding task as an example, although VLMs can detect and locate objects in the image, the accuracy can not be as good as the results obtained by object detectors. So in my personal view, LLMs/VLMs are not good at fine-grained spatial tasks at the current time point.

Can you provide some insight about the superior performance? Thanks!

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

About the problem formulation and evaluation results #7

Metadata

Assignees

Labels

Projects

Milestone

Relationships

Development

About the problem formulation and evaluation results #7

Description

Metadata

Metadata

Assignees

Labels

Projects

Milestone

Relationships

Development

Issue actions