# Error when using Langchain WebResearchRetriever – RuntimeError: asyncio.run() cannot be called from a running event loop

**URL:** https://community.openai.com/t/error-when-using-langchain-webresearchretriever-runtimeerror-asyncio-run-cannot-be-called-from-a-running-event-loop/341969
**Category:** API
**Tags:** langchain, python
**Created:** [August 31, 2023, 12:59am UTC](https://community.openai.com/t/error-when-using-langchain-webresearchretriever-runtimeerror-asyncio-run-cannot-be-called-from-a-running-event-loop/341969 "2023-08-31T00:59:29Z")
**Posts on this page:** 9
**Page:** 1

<div class="post-metadata">

### Author: ![arnimk](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/arnimk/32/188804_2.png) [@arnimk](https://community.openai.com/u/arnimk)
#### Post date: [August 31, 2023, 12:59am UTC](https://community.openai.com/t/error-when-using-langchain-webresearchretriever-runtimeerror-asyncio-run-cannot-be-called-from-a-running-event-loop/341969/1 "2023-08-31T00:59:30Z")

</div>

I’ve been attempting to use WebResearchRetriever from Langchain in Python, and I’m running a segment of code that works for other people, but I keep getting this error:

RuntimeError: asyncio.run() cannot be called from a running event loop

I think the issue may be with my computer and not with the code itself, but here’s the code:

from langchain.retrievers.web\_research import WebResearchRetriever  
import os  
from langchain.vectorstores import Chroma  
from langchain.embeddings import OpenAIEmbeddings  
from langchain.chat\_models.openai import ChatOpenAI  
from langchain.utilities import GoogleSearchAPIWrapper

os.environ[“OPENAI\_API\_KEY”] = ‘my\_key’

vectorstore = Chroma(embedding\_function=OpenAIEmbeddings(),persist\_directory=“./chroma\_db\_oai”)

llm = ChatOpenAI(temperature=0)

os.environ[“GOOGLE\_CSE\_ID”] = “my\_key”  
os.environ[“GOOGLE\_API\_KEY”] = “my\_key”  
search = GoogleSearchAPIWrapper()

web\_research\_retriever = WebResearchRetriever.from\_llm(  
vectorstore=vectorstore,  
llm=llm,  
search=search,  
)

from langchain.chains import RetrievalQAWithSourcesChain  
user\_input = “How do LLM Powered Autonomous Agents work?”  
qa\_chain = RetrievalQAWithSourcesChain.from\_chain\_type(llm,retriever=web\_research\_retriever)  
result = qa\_chain({“question”: user\_input})  
print (result)

Can anyone help me resolve this error? Any help would be much appreciated.

---

<div class="post-metadata">

### Author: ![novaphil](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/novaphil/32/130934_2.png) [@novaphil](https://community.openai.com/u/novaphil)
#### Post date: [August 31, 2023, 3:59pm UTC](https://community.openai.com/t/error-when-using-langchain-webresearchretriever-runtimeerror-asyncio-run-cannot-be-called-from-a-running-event-loop/341969/2 "2023-08-31T15:59:08Z")

</div>

You’ll probably have better luck asking on the LangChain forums since the issue is with their code.

---

<div class="post-metadata">

### Author: ![hydeta](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/hydeta/32/189964_2.png) [@hydeta](https://community.openai.com/u/hydeta)
#### Post date: [September 1, 2023, 1:12am UTC](https://community.openai.com/t/error-when-using-langchain-webresearchretriever-runtimeerror-asyncio-run-cannot-be-called-from-a-running-event-loop/341969/3 "2023-09-01T01:12:55Z")

</div>

Are you running this in Jupyter? If so they already run an event loop in the background.

---

<div class="post-metadata">

### Author: ![arnimk](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/arnimk/32/188804_2.png) [@arnimk](https://community.openai.com/u/arnimk)
#### Post date: [September 2, 2023, 11:35pm UTC](https://community.openai.com/t/error-when-using-langchain-webresearchretriever-runtimeerror-asyncio-run-cannot-be-called-from-a-running-event-loop/341969/4 "2023-09-02T23:35:47Z")

</div>

I’ve tried it in Jupyter and on Google Colab. Do you know what I should use instead so that there isn’t already an event loop?

---

<div class="post-metadata">

### Author: ![BabellDev](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/babelldev/32/448374_2.png) [@BabellDev](https://community.openai.com/u/BabellDev)
#### Post date: [September 7, 2023, 2:31pm UTC](https://community.openai.com/t/error-when-using-langchain-webresearchretriever-runtimeerror-asyncio-run-cannot-be-called-from-a-running-event-loop/341969/5 "2023-09-07T14:31:24Z")

</div>

I had this same issue in Jupyter. You should be able to solve it by tweaking your code similar to this:

# Make sure nest\_asyncio is installed  
!pip install nest\_asyncio

# Allow nested asyncio loops  
import nest\_asyncio  
nest\_asyncio.apply()

# Move your use of qa\_chain into an async function  
async def main():  
result = qa\_chain({“question”: user\_input})  
print (result)

# Run your async function in the existing event loop  
loop = asyncio.get\_event\_loop()  
loop.run\_until\_complete(main())

---

<div class="post-metadata">

### Author: ![arnimk](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/arnimk/32/188804_2.png) [@arnimk](https://community.openai.com/u/arnimk)
#### Post date: [September 12, 2023, 10:03pm UTC](https://community.openai.com/t/error-when-using-langchain-webresearchretriever-runtimeerror-asyncio-run-cannot-be-called-from-a-running-event-loop/341969/6 "2023-09-12T22:03:04Z")

</div>

Thank you so much! The error is gone. However, I now have another error which I have no clue how to fix. Any idea how to fix this?

`ClientConnectorCertificateError: Cannot connect to host python.langchain.com:443 ssl:True [SSLCertVerificationError: (1, '[SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: unable to get local issuer certificate (_ssl.c:1002)')]`

---

<div class="post-metadata">

### Author: ![BabellDev](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/babelldev/32/448374_2.png) [@BabellDev](https://community.openai.com/u/BabellDev)
#### Post date: [September 13, 2023, 2:32pm UTC](https://community.openai.com/t/error-when-using-langchain-webresearchretriever-runtimeerror-asyncio-run-cannot-be-called-from-a-running-event-loop/341969/7 "2023-09-13T14:32:30Z")

</div>

Actually I’ve been getting that same SSL error while utilizing `WebResearchRetriever`. From what I’ve read, I think the _true_ fix involves tweaking local Python SSL certificate config/handling, but things I’ve tried haven’t worked.

However, I did come up with a monkey-patch you can use to modify the behavior of `WebResearchRetriever._get_relevant_documents` to disable SSL cert verification. This isn’t ideal from a security point-of-view, but it works. This updated logic also adds a check for empty docs list before adding to the vector database, which solves a tuple exception when no docs could be decoded (if they are all pdfs, for example). I’ve added comments with my name `# (BabellDev)` so you can see the changes. Otherwise the code is the same as the current version on github.

Add this function declaration somewhere before your existing code, and then call it before you call  
`WebResearchRetriever.from_llm`.

```
def patch_web_research_retriever():
    import logging
    from typing import List
    from langchain.retrievers.web_research import WebResearchRetriever
    from langchain.callbacks.manager import CallbackManagerForRetrieverRun
    from langchain.document_loaders import AsyncHtmlLoader
    from langchain.document_transformers import Html2TextTransformer
    from langchain.schema import Document

    logger = logging.getLogger( __name__ )

    def _patched_get_relevant_documents(
        self,
        query: str,
        *,
        run_manager: CallbackManagerForRetrieverRun,
    ) -> List[Document]:

        # Get search questions
        logger.info("Generating questions for Google Search ...")
        result = self.llm_chain({"question": query})
        logger.info(f"Questions for Google Search (raw): {result}")
        questions = getattr(result["text"], "lines", [])
        logger.info(f"Questions for Google Search: {questions}")

        # Get urls
        logger.info("Searching for relevant urls...")
        urls_to_look = []
        for query in questions:
            # Google search
            search_results = self.search_tool(query, self.num_search_results)
            logger.info("Searching for relevant urls...")
            logger.info(f"Search results: {search_results}")
            for res in search_results:
                if res.get("link", None):
                    urls_to_look.append(res["link"])

        # Relevant urls
        urls = set(urls_to_look)

        # Check for any new urls that we have not processed
        new_urls = list(urls.difference(self.url_database))

        logger.info(f"New URLs to load: {new_urls}")
        # Load, split, and add new urls to vectorstore
        if new_urls:

            # (BabellDev) changed verify_ssl to False
            loader = AsyncHtmlLoader(new_urls, verify_ssl=False)
            html2text = Html2TextTransformer()
            logger.info("Indexing new urls...")
            docs = loader.load()
            docs = list(html2text.transform_documents(docs))
            docs = self.text_splitter.split_documents(docs)

            # (BabellDev) do not add if docs is empty (avoid tuple error)
            if docs is not None and len(docs) > 0:
                self.vectorstore.add_documents(docs)

            self.url_database.extend(new_urls)

        # Search for relevant splits
        logger.info("Grabbing most relevant splits from urls...")
        docs = []
        for query in questions:
            docs.extend(self.vectorstore.similarity_search(query))

        # Get unique docs
        unique_documents_dict = {
            (doc.page_content, tuple(sorted(doc.metadata.items()))): doc for doc in docs
        }
        unique_documents = list(unique_documents_dict.values())
        return unique_documents

    WebResearchRetriever._get_relevant_documents = _patched_get_relevant_documents

```

If you don’t like monkey-patching, you could derive your own class from `WebResearchRetriever` and override the `_get_relevant_documents` method.

Hope it helps!

---

<div class="post-metadata">

### Author: ![addarcher](https://avatars.discourse-cdn.com/v4/letter/a/8edcca/32.png) [@addarcher](https://community.openai.com/u/addarcher)
#### Post date: [October 3, 2023, 2:50pm UTC](https://community.openai.com/t/error-when-using-langchain-webresearchretriever-runtimeerror-asyncio-run-cannot-be-called-from-a-running-event-loop/341969/8 "2023-10-03T14:50:18Z")

</div>

had same issue. I think the solution is to adjust the async\_html.py in \Lib\site-packages\langchain\document\_loaders.

first add these two:  
import ssl  
import certifi

then change the code block at the bottom to:

ssl\_context = ssl.create\_default\_context(cafile=certifi.where())  
conn = aiohttp.TCPConnector(ssl=ssl\_context)  
async with aiohttp.ClientSession(connector=conn) as session:

---

<div class="post-metadata">

### Author: ![BabellDev](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/babelldev/32/448374_2.png) [@BabellDev](https://community.openai.com/u/BabellDev)
#### Post date: [October 3, 2023, 3:42pm UTC](https://community.openai.com/t/error-when-using-langchain-webresearchretriever-runtimeerror-asyncio-run-cannot-be-called-from-a-running-event-loop/341969/9 "2023-10-03T15:42:58Z")

</div>

FYI. Regarding the original question, it looks like this has been fixed in the latest version of langchain. If you take a look at `load()` in `async_html` it is now handling an already-running loop:

> <https://github.com/langchain-ai/langchain/blob/master/libs/langchain/langchain/document_loaders/async_html.py>
