✈ TravelAgent Demo

Research brief 03 · Honest strengths + weaknesses · Source: 03-llm-agent-test.md

What we tested

Prefilled request (from test-location.md): 9-day Kyoto + Tokyo/Hakone for 2, mid-range, culture + food + calm pace, avoid crowds, private onsen once, tea ceremony, Nara/Uji day trip. Day-by-day plan, hotels, transport, bookings, budget, pitfalls.

Demo app (app.py) sends this to Muse Spark 1.3 — or uses curated fallback without an API key. See the frozen output: Agent Portal Demo.

What the LLM does well TODAY

Genuinely useful for a beginner agent: 70% of the research in seconds.

Where it falls short vs. a seasoned human

  1. No live truth: doesn't know what's sold out or price-changed today. Hallucinates times/prices if pushed.
  2. Generic picks: "top 10" hotels/temples, not the quiet machiya or the guide who's great with kids.
  3. No relationships: can't call the ryokan, secure the upgrade, or fix a cancellation at 9pm.
  4. Weak trade-offs: "both are great" instead of "skip Kinkaku-ji at noon, do it at open + add Uji."
  5. No accountability: if the plan fails, there's no one to call. Trust drops fast on high-spend trips.
  6. Shallow season nuance: knows "Golden Week is busy" but not the hour-by-hour temple strategy.

Scorecard (1–5)

LLM (Muse Spark 1.3)Seasoned agent
Speed52
Breadth of ideas53
Accuracy (live prices/availability)25
Personalization35
Insider access / upgrades15
Crisis help15
Trust / accountability25

Takeaway — build for tomorrow, not just today