Presets & Research Archetypes
🎓 Academic
ArXiv, peer review & whitepapers
💬 Discussions
Reddit, HackerNews, forum consensus
🌐 Web (Deep)
Broad indexed industry reports & news
🔢 Wolfram/Code
Mathematical proofs, code & data
Entities Extracted: 4
Recommended Route: Academic & Technical
🎯 Recommended Primary Boolean Query
"enterprise LLM deployment" AND ("on-premise" OR "vLLM" OR "TGI" OR "Ollama") AND ("latency" OR "cost per token") filetype:pdf site:arxiv.org OR site:openreview.net
⚡ Strict Boolean Syntax String
"enterprise LLM deployment" AND ("on-premise" OR "vLLM" OR "TGI" OR "Ollama") AND ("latency" OR "cost per token") site:arxiv.org OR site:openreview.net
🔀 Multi-Hop Sub-Query Vectors
2 vectors
What are the total cost of ownership (TCO) break-even thresholds for hosted APIs vs self-hosted 70B models?
Benchmarks comparing vLLM throughput against OpenAI batch API in production workloads
🛡️ Search Boundary & Domain Filters
arxiv.org
semianalysis.com
anyscale.com
github.com
📋 Citation & Evidence Verification Directive
Request explicit quantitative tables with token throughput (tok/s) and GPU hardware specs.