tlooto: An Evidence-Grounded Agentic Workspace for Scholarly Synthesis and Manuscript Development
tlooto Research Team
Abstract
Background. The scholarly literature grows by roughly 4% a year, and a single systematic review takes more than a year of expert effort. Large language models (LLMs) promise relief, yet frontier models fabricate most of the scientific citations they produce, and even retrieval-based commercial tools hallucinate in up to a third of queries.
Objective. We present tlooto, a research workspace that answers research questions from the literature and develops the answer into a manuscript, designed so that every citation points to a scholarly record retrieved for the question.
System. tlooto coordinates planning, scholarly retrieval, literature screening, evidence synthesis and editing in a single agentic workflow. Each stage is shown to the researcher as it happens, and synthesis draws its citations only from papers that passed screening.
Capabilities. Beyond question answering, tlooto provides a paper-format editor with live citations, tables, figures and equations in five citation styles, direct paper search, a personal library and project memory, and journal recommendation, covering the research lifecycle from ideation to submission.
Significance. tlooto makes citation reliability a property of how the answer is produced rather than something the researcher must verify afterwards, and brings the emerging class of research agents into a form designed for everyday scholarly work.
Keywords: research agents; retrieval-augmented generation; citation grounding; evidence synthesis; literature review automation; scholarly writing
1. Introduction
1.1 The scale of the literature
Scientific output has grown continuously for more than a century. Bornmann et al. estimate current growth at about 4.10% a year, a doubling time of roughly 17 years [1]. In clinical medicine the pressure is sharper still: already in 2010, some 75 trials and 11 systematic reviews were published every day [2] (Figure 1). Synthesising this evidence by hand is slow and costly; systematic reviews registered in PROSPERO took a mean of 67.3 weeks from registration to publication [3].

Figure 1. Growth of the scientific literature. (a) Publication output modelled at the reported annual growth rate of 4.1%, a doubling time of about 17 years. (b) Trials and systematic reviews published per day in medicine in 2010.
1.2 AI enters research practice
Researchers have responded by adopting AI. In a Nature survey of more than 1,600 scientists, a majority expected AI tools to become very important or essential to their field [4], and an analysis of the published record found LLM-modified content in up to 17.5% of recent computer-science papers [5]. Research agents have followed rapidly: they synthesise scientific knowledge [13] [14] [15], draft survey articles [16], generate research ideas that expert reviewers rate as more novel than human ideas [19], and propose and test hypotheses [17] [18].
1.3 The reliability gap
General-purpose LLMs, however, generate fluent statements that no source supports [6], and citations are where this failure is most visible and most costly. In the OpenScholar evaluation, GPT-4o fabricated 78–90% of the scientific citations it produced [15]. Retrieval alone does not close the gap: leading commercial research tools built on retrieval still hallucinated in 17–33% of queries [7] (Figure 2). For researchers, a citation that cannot be trusted is worse than no citation at all.

Figure 2. Reported rates of fabricated or unsupported citations in recent evaluations of AI research assistance.
1.4 Contributions
This paper describes tlooto, a system built so that researchers can rely on its citations. Our contributions are:
- Evidence-grounded synthesis, in which a report cites only papers retrieved and screened for the question at hand (Section 3).
- An agentic research workflow that plans, retrieves, screens, synthesises and edits in one continuous process (Section 4).
- A transparent Work process that shows the researcher each plan, query, screening decision and intermediate result in real time (Section 4.4).
- An integrated manuscript environment that carries cited evidence into a full editor and on to journal selection (Section 5).
2. Related work
2.1 Retrieval-augmented and citation-aware generation
Retrieval-augmented generation (RAG) conditions a language model on documents fetched at inference time [8] and has become the dominant approach to knowledge-intensive tasks [9]. Gao et al. showed that models can be prompted and evaluated to attach citations to the passages that support each statement [10]. Agentic formulations interleave reasoning with tool use [11], building on chain-of-thought prompting [12].
2.2 Research agents for scientific literature
PaperQA introduced an agent that retrieves full-text papers to answer scientific questions [13], and PaperQA2 reported synthesis of scientific knowledge that exceeded expert performance on its benchmarks [14]. OpenScholar combined a large open-access corpus with a retrieval-augmented model and self-feedback to answer literature questions with citations [15]. AutoSurvey drafts full survey articles [16], while The AI Scientist [17] and AI co-scientist [18] extend agents to idea generation, experimentation and hypothesis testing.
2.3 Automation of evidence synthesis
Machine learning has long supported the screening stage of systematic reviews [20]. Active-learning tools such as ASReview substantially reduce the number of records that reviewers must read [21], and LLM-based screening of clinical reviews has achieved high agreement with human decisions [22]. Reporting standards such as PRISMA 2020 describe the stages that such tools accelerate [23].
2.4 Positioning
Table 1 contrasts tlooto with the tools researchers most often use today. tlooto combines agentic retrieval and screening, citation-grounded synthesis, a visible reasoning process and a full writing environment in a single workspace built for daily research.
Table 1. Capability comparison with tools researchers commonly use.
| Capability |
General AI chatbot |
Academic search engine |
tlooto |
| Answers a research question in prose |
Yes |
No |
Yes |
| Claims grounded in retrieved papers |
No |
Not applicable |
Yes |
| Citations drawn from retrieved records |
No |
Not applicable |
Yes |
| Screening against explicit criteria |
No |
No |
Yes |
| Step-by-step process visible to the user |
No |
No |
Yes (Work process) |
| Manuscript editor with live citations |
No |
No |
Yes |
| Five citation styles, PDF and DOCX export |
No |
Export only |
Yes |
| Library, projects and reusable evidence |
No |
Library only |
Yes |
| Journal recommendation |
No |
No |
Yes |
3. Evidence-grounded synthesis
Let $q$ denote a research question. tlooto issues a set of search queries $Q(q)$ and collects the retrieved records
$$R(q) = \bigcup_{k \in Q(q)} \operatorname{retrieve}(k)$$
Screening applies inclusion criteria $\phi$ derived from the question and keeps the evidence set
$$E(q) = {, r \in R(q) : \phi(r, q) = 1 ,}$$
The report $A$ is written as a sequence of statements $s_1, \dots, s_n$, each carrying a citation set $c(s_i)$. tlooto is designed around the grounding principle
$$\bigcup_{i=1}^{n} c(s_i) ;\subseteq; E(q) ;\subseteq; R(q)$$
so that the reference list of a report is built from the same scholarly records the answer was written from. Because every record in $R(q)$ comes from a scholarly index with its DOI and bibliographic metadata, the citations in a tlooto report lead to papers that exist and that the researcher can open.
4. The tlooto workflow

Figure 3. The tlooto workflow from research question to manuscript. Screening and synthesis are highlighted; the Work process makes every stage visible to the researcher.
4.1 Planning and scholarly retrieval
tlooto first interprets the question together with any scope, discipline and inclusion criteria the researcher supplies, and breaks it into complementary search intents. Rather than issuing a single query, it runs several reformulated queries per intent against a scholarly index and merges the results. Every retrieved record carries its authors, venue, year, abstract and DOI.
4.2 Literature screening
Each candidate is assessed against explicit inclusion criteria drawn from the question, mirroring the screening stage of a systematic review [20] [21] [22] [23]. The papers that pass are ranked and passed to synthesis as a focused evidence set rather than a raw list of search results.
4.3 Evidence synthesis
The report is written from the screened papers. Each statement carries a numbered citation, and the reference list is generated from the same records, so the body and the bibliography always agree. Follow-up questions build on the evidence already gathered, so a line of inquiry deepens rather than restarts.
4.4 The Work process
Every plan, query, screening decision and intermediate result is streamed to the researcher in the Work process panel. Researchers can see what was searched, which papers were kept and how the answer was assembled, turning the system's reasoning into a record they can follow and cite in their own methods.
4.5 Research Editor
With one action the report opens in the Research Editor, a paper-format editor that supports headings, tables, figures, equations and live citation objects. Citations renumber automatically as text moves, render in APA, MLA, Harvard, Chicago or ISO 690, and export with the manuscript to PDF or DOCX. Researchers can ask tlooto to revise, extend or restructure sections while every citation stays attached to its source.
Table 2. Stages of the tlooto workflow and what the researcher receives at each.
| Stage |
What tlooto does |
What the researcher sees |
| Planning |
Interprets the question, scope and criteria |
Search intents in the Work process |
| Retrieval |
Runs multiple queries on a scholarly index |
Retrieved papers with DOI and metadata |
| Screening |
Applies inclusion criteria and ranks papers |
Which papers were kept and why |
| Synthesis |
Writes a structured report from the kept papers |
Cited report with a matching reference list |
| Editing |
Develops the report into a manuscript |
Paper-format draft with live citations |
5. Coverage of the research lifecycle

Figure 4. Research stages supported by a general AI chatbot, an academic search engine and tlooto.
tlooto is designed as one workspace for the whole project rather than a single-purpose tool (Figure 4, Table 3). Questions, saved papers and uploaded files accumulate in projects and the library, so later questions build on earlier evidence instead of starting a new search.
Table 3. Research tasks supported by tlooto.
| Research task |
What tlooto delivers |
Where |
| Framing a research question |
Candidate questions and theoretical perspectives with supporting literature |
Report |
| Literature review |
Map of perspectives, consensus and debates, fully cited |
Report, Paper Search |
| Gap analysis |
Underexplored areas with the evidence that defines them |
Report |
| Methodology |
Designs, measures and analyses tied to studies that used them |
Report |
| Evidence management |
Saved papers, uploaded files and reusable evidence |
My Library, Projects |
| Manuscript drafting |
Paper-format drafting with live citations, tables, figures and equations |
Research Editor |
| Referencing |
Five citation styles, automatic renumbering, PDF and DOCX export |
Research Editor |
| Venue selection |
Journal recommendations matched to the manuscript |
AI Journal Finder |
Table 4. Example questions and the report structure tlooto produces.
| Question type |
Example question |
Report structure |
| Research question |
How is generative AI changing knowledge work in professional organisations? |
Perspectives, candidate questions, supporting studies |
| Research gap |
What remains underexplored in research on AI adoption in organisations? |
Gaps, why each matters, how to address it |
| Literature review |
How has platform ecosystem governance research evolved over five years? |
Themes, agreements, debates, trajectory |
| Evidence |
Does AI investment improve firm performance? |
Supporting and conflicting findings, moderators |
| Methodology |
Does psychological safety mediate AI adoption and job satisfaction? |
Design, measures, sampling, analysis plan |
6. Discussion
The central design decision in tlooto is to treat citation grounding as part of how an answer is produced rather than as a check applied afterwards. Evaluations of retrieval-based commercial tools show that retrieval by itself still leaves substantial error [7]; tlooto adds explicit screening and builds each report's references from the screened papers, and it exposes the whole process so researchers can see how every conclusion was reached.
The second decision is integration. Research agents have shown that machines can synthesise literature at expert level [14] [15] and generate strong research ideas [19], yet a researcher's work does not end with an answer. tlooto carries the evidence directly into the manuscript, keeps citations live through revision and supports journal selection, so the time saved in search and screening is not lost again in formatting and referencing.
Together these choices make tlooto a practical research partner for the workload described in Section 1: one that reads widely, cites faithfully and writes in the form that scholarship requires.
7. Conclusion
The literature now grows faster than any researcher can read, and general-purpose AI cannot yet be trusted with references. tlooto addresses both problems with an agentic workspace that retrieves and screens before it writes, builds every reference list from the papers it kept, shows its work step by step, and develops the result into a manuscript within a single workspace.