Sure To Fall

What does the life cycle of a scientific paper look like?

It stands to reason that after a paper is published, people download and read the paper and then if it generates sufficient interest, it will begin to be cited. At some point these citations will peak and the interest will die away as the work gets superseded or the field moves on. So each paper has a useful lifespan. When does the average paper start to accumulate citations, when do they peak and when do they die away?

Citation behaviours are known to be very field-specific. So to narrow things down, I focussed on cell biology and in one area “clathrin-mediated endocytosis” in particular. It’s an area that I’ve published in – of course this stuff is driven by self-interest. I downloaded data for 1000 papers from Web of Science that had accumulated the most citations. Reviews were excluded, as I assume their citation patterns are different from primary literature. The idea was just to take a large sample of papers on a topic. The data are pretty good, but there are some errors (see below).

Number-crunching (feel free to skip this bit): I imported the data into IgorPro making a 1D wave for each record (paper). I deleted the last point corresponding to cites in 2014 (the year is not complete). I aligned all records so that year of publication was 0. Next, the citations were normalised to the maximum number achieved in the peak year. This allows us to look at the lifecycle in a sensible way. Next I took out records to papers less than 6 years old as I reasoned these would have not have completed their lifecycle and could contaminate the analysis (it turned out to make little difference). The lifecycles were plotted and averaged. I also wrote a quick function to pull out the peak year for citations post hoc.

So what did it show?

Citations to a paper go up and go down, as expected (top left). When cumulative citations are plotted most of the articles have an initial burst and then level off. The exception are ~8 articles that continue to rise linearly (top right). On average a paper generates its peak citations three years after publication (box plot). The fall after this peak period is pretty linear and it’s apparently all over somewhere >15 years after publication (bottom left). To look at the decline in more detail I aligned the papers so that year 0 was the year of peak citations. The average now loses almost 40% of those peak citations in the following year and then declines steadily (bottom right).

Edit: The dreaded Impact Factor calculation takes the citations to articles published in the preceding 2 years and divides by the number of citable items in that period. This means that each paper only contributes to the Impact Factor in years 1 and 2. This is before the average paper reaches its peak citation period. Thanks to David Stephens (@david_s_bristol) for pointing this out. The alternative 5 year Impact Factor gets around this limitation.

Perhaps lifecycle is the wrong term: papers in this dataset don’t actually ‘die’, i.e. go to 0 citations. There is always a chance that a paper will pick up the odd citation. Papers published 15 years ago are still clocking 20% of their peak citations. Looking at papers cited at lower rates would be informative here.

Two other weaknesses that affect precision is that 1) a year is a long time and 2) publication is subject to long lag times. The analysis would be improved by categorising the records based on the month-year when the paper was published and the month-year when each citation comes in. Papers published in January in one year probably have a different peak than those published in December of the same year, but this is lost when looking at year alone. Secondly, due to publication lag, it is impossible to know when the peak period of influence for a paper truly is.
MisCytesProblems in the dataset. Some reviews remained despite being supposedly excluded, i.e. they are not properly tagged in the database. Also, some records have citations from years before the article was published! The numbers of citations are small enough to not worry for this analysis, but it makes you wonder about how accurate the whole dataset is. I’ve written before about how complete citation data may or may not be. These sorts of things are a concern for all of us who are judged by these things for hiring and promotion decisions.

The post title is taken from ‘Sure To Fall’ by The Beatles, recorded during The Decca Sessions.

Outer Limits

This post is about a paper that was recently published. It was the result of a nice collaboration between me and Francisco López-Murcia and Artur Llobet in Barcelona.

The paper in a nutshell
The availability of clathrin sets a limit for presynaptic function

Clathrin is a three legged protein that forms a cage around membranes during endoctosis. One site of intense clathrin-mediated endocytosis (CME) is the presynaptic terminal. Here, synaptic vesicles need to be recaptured after fusion and CME is the main route of retrieval. Clathrin is highly abundant in all cells and it is generally thought of as limitless for the formation of multiple clathrin-coated structures. Is this really true? In a neuron where there is a lot of endocytic activity, maybe the limits are tested?
It is known that strong stimulation of neurons causes synaptic depression – a form of reversible synaptic plasticity where the neuron can only evoke a weak postsynaptic response afterwards. Is depression a vesicle supply problem?

What did we find?
We showed that clathrin availability drops during stimulation that evokes depression. The drop in availability is due to clathrin forming vesicles and moving away from the synapse. We mimicked this by RNAi, dropping the clathrin levels and looking at synaptic responses. We found that when the clathrin levels drop, synaptic responses become very small. We noticed that fewer vesicles are able to be formed and those that do form are smaller. Interestingly, the amount of neurotransmitter (acetylcholine) in the vesicles was much less than the volume of the vesicles as measured by electron microscopy. This suggests there is an additional sorting problem in cells with lower clathrin levels.

Killer experiment
A third reviewer was called in (due to a split decision between Reviewers 1 and 2). He/she asked a killer question: all of our data could be due to an off-target effect of RNAi, could we do a rescue experiment? We spent many weeks to get the rescue experiment to work, but a second viral infection was too much for the cells and engineering a virus to express clathrin was very difficult. The referee also said: if clathrin levels set a limit for synaptic function, why don’t you just express more clathrin? Well, we would if we could! But this gave us an idea… why don’t we just put clathrin in the pipette and let it diffuse out to the synapses and rescue the RNAi phenotype over time? We did it – and to our surprise – it worked! The neurons went from an inhibited state to wild-type function in about 20 min. We then realised we could use the same method on normal neurons to boost clathrin levels at the synapse and protect against synaptic depression. This also worked! These killer experiments were a great addition to the paper and are a good example of peer review improving the paper.

Fran and Artur did almost all the experimental work. I did a bit of molecular biology and clathrin purification. Artur and I wrote the paper and put the figures together – lots of skype and dropbox activity.
Artur is a physiologist and his lab like to tackle problems that are experimentally very challenging – work that my lab wouldn’t dare to do – he’s the perfect collaborator. I have known Artur for years. We were postdocs in the same lab at the LMB in the early 2000s. We tried a collaborative project to inhibit dynamin function in adrenal chromaffin cells at that time, but it didn’t work out. We have stayed in touch and this is our first paper together. The situation in Spain for scientific research is currently very bad and it deteriorated while the project was ongoing. This has been very sad to hear about, but fortunately we were able to finish this project and we hope to work together more in the future.

We were on the cover!
25.coverNow the scientific literature is online, this doesn’t mean so much anymore, but they picked our picture for the cover. It is a single cell microculture expressing GFP that was stained for synaptic markers and clathrin. I changed the channels around for artistic effect.

What else?
J Neurosci is slightly different to other journals that I’ve published in recently (my only other J Neurosci paper was published in 2002). For the following reasons:

  1. No supplementary information. The journal did away with this years ago to re-introduce some sanity in the peer review process. This didn’t affect our paper very much. We had a movie of clathrin movement that would have gone into the SI at another journal, but we simply removed it here.
  2. ORCIDs for authors are published with the paper. This gives the reader access to all your professional information and distinguishes authors with similar names. I think this is a good idea.
  3. Submission fee. All manuscripts are subject to a submission fee. I believe this is to defray the costs of editorial work. I think this makes sense, although I’m not sure how I would feel if our paper had been rejected.


López-Murcia, F.J., Royle, S.J. & Llobet, A. (2014) Presynaptic clathrin levels are a limiting factor for synaptic transmission J. Neurosci., 34: 8618-8629. doi: 10.1523/JNEUROSCI.5081-13.2014

Pubmed | Paper

The post title is taken from “Outer Limits” a 7″ Single by Sleep ∞ Over released in 2010.

All This And More

I was looking at the latest issue of Cell and marvelling at how many authors there are on each paper. It’s no secret that the raison d’être of Cell is to publish the “last word” on a topic (although whether it fulfils that objective is debatable). Definitive work needs to be comprehensive. So it follows that this means lots of techniques and ergo lots of authors. This means it is even more impressive when a dual author paper turns up in the table of contents for Cell. Anyway, I got to thinking: has it always been the case that Cell papers have lots of authors and if not, when did that change?

I downloaded the data for all articles published by Cell (and for comparison, J Cell Biol) from Scopus. The records required a bit of cleaning. For example, SnapShot papers needed to be removed and also the odd obituary etc. had been misclassified as an article. These could be quickly removed. I then went back through and filtered out ‘articles’ that were less than three pages as I think it is not possible for a paper to be two pages or fewer in length. The data could be loaded into IgorPro and boxplots generated per year to show how author number varied over time. Reviews that are misclassified as Articles will still be in the dataset, but I figured these would be minimal.

Authors1First off: Yes, there are more authors on average for a Cell paper versus a J Cell Biol paper. What is interesting is that both journals had similar numbers of authors when Cell was born (1974) and they crept up together until the early 2000s, when the number of Cell authors kept increasing, or JCell Biol flattened off, whichever way you look at it.

I think the overall trend to more authors is because understanding biology has increasingly required multiple approaches and the bar for evidence seems to be getting higher over time. The initial creep to more authors (1974-2000) might be due to a cultural change where people (technicians/students/women) began to get proper credit for their contributions. However, this doesn’t explain the divergence between J Cell Biol and Cell in recent years. One possibility is Cell takes more non-cell biology papers and that these papers necessarily have more authors. For example, the polar bear genome was published in Cell (29 authors), and this sort of paper would not appear in J Cell Biol. Another possibility is that J Cell Biol has a shorter and stricter revision procedure, which means that multiple rounds of revision, collecting new techniques and new authors is more limited than it is at Cell. Any other ideas?

AuthorI also quickly checked whether more authors means more citations, but found no evidence for such a relationship. For papers published in the years 2000-2004, the median citation number for papers with 1-10 authors was pretty constant for J Cell Biol. For Cell, these data mere more noisy. Three-author papers tended to be cited a bit more than those with two authors, but then four author papers were also lower.

The number of authors on papers from our lab ranges from 2-9 and median is 3.5. This would put an average paper from our lab in the bottom quartile for JCB and in the lower 10% for Cell in 2013. Ironically, our 9 author paper (an outlier) was published in J Cell Biol. Maybe we need to get more authors on our papers before we can start troubling Cell with our manuscripts…

The Post title is taken from ‘All This and More’ by The Wedding Present from their LP George Best.

Into The Great Wide Open

We have a new paper out! You can read it here.

I thought I would write a post on how this paper came to be and also about our first proper experience with preprinting.

Title of the paper: Non-specificity of Pitstop 2 in clathrin-mediated endocytosis.

In a nutshell: we show that Pitstop 2, a supposedly selective clathrin inhibitor acts in a non-specific way to inhibit endocytosis.

Authors: Anna Willox, who was a postdoc in the lab from 2008-2012, did the flow cytometry measurements and Yasmina Sahraoui who was a summer student in my lab, did the binding experiments. And me.

Background: The description of “pitstops” – small molecules that inhibit clathrin-mediated endocytosis – back in 2011 in Cell was heralded as a major step-forward in cell biology. And it really would be a breakthrough if we had ways to selectively switch off clathrin-mediated endocytosis. Lots of nasty things gain entry into cells by hijacking this pathway, including viruses such as HIV and so if we could stop viral entry this could prevent cellular infection. Plus, these reagents would be really handy in the lab for cell biologists.

The rationale for designing the pitstop inhibitors was that they should block the interaction between clathrin and adaptor proteins. Adaptors are the proteins that recognise the membrane and cargo to be internalised – clathrin itself cannot do this. So if we can stop clathrin from binding adaptors there should be no internalisation – job done! Now, in 2000 or so, we thought that clathrin binds to adaptors via a single site on its N-terminal domain. This information was used in the drug screen that identified pitstops. The problem is that, since 2000, we have found that there are four sites on the N-terminal domain of clathrin that can each mediate endocytosis. So blocking one of these sites with a drug, would do nothing. Despite this, pitstop compounds, which were shown to have a selectivity for one site on the N-terminal domain of clathrin, blocked endocytosis. People in the field scratched their hands at how this is possible.

A damning paper was published in 2012 from Julie Donaldson’s lab showing that pitstops inhibit clathrin-independent endocytosis as well as clathrin-mediated endocytosis. Apparently, the compounds affect the plasma membrane and so all internalisation is inhibited. Many people thought this was the last that we would hear about these compounds. After all, these drugs need to be highly selective to be any use in the lab let alone in the clinic.

Our work: we had our own negative results using these compounds, sitting on our server, unpublished. Back in February 2011, while the Pitstop paper was under revision, the authors of that study sent some of these compounds to us in the hope that we could use these compounds to study clathrin on the mitotic spindle. The drugs did not affect clathrin binding to the spindle (although they probably should have done) and this prompted us to check whether the compounds were working – they had been shipped all the way from Australia so maybe something had gone wrong. We tested for inhibition of clathrin-mediated endocytosis and they worked really well.

At the time we were testing the function of each of the four interaction sites on clathrin in endocytosis, so we added Pitstop 2 to our experiments to test for specificity. We found  that Pitstop 2 inhibits clathrin-mediated endocytosis even when the site where Pitstops are supposed to bind, has been mutated! The picture shows that the compound (pink) binds where sequences from adaptors can bind. Mutation of this site doesn’t affect endocytosis, because clathrin can use any three of the other four sites. Yet Pitstop blocks endocytosis mediated by this mutant, so it must act elsewhere, non-specifically.

So the compounds were not as specific as claimed, but what could we do with this information? There didn’t seem enough to publish and I didn’t want people in the lab working on this as it would take time and energy away from other projects. Especially when debunking other people’s work is such a thankless task (why this is the case, is for another post). The Dutta & Donaldson paper then came out, which was far more extensive than our results and so we moved on.

What changed?

A few things prompted me to write this work up. Not least, Yasmina had since shown that our mutations were sufficient to prevent AP-2 binding to clathrin. This result filled a hole in our work. These things were:

  1. People continuing to use pitstops in published work, without acknowledging that they may act non-specifically. The turning point was this paper, which was critical of the Dutta & Donaldson work.
  2. People outside of the field using these compounds without realising their drawbacks.
  3. AbCam selling this compound and the thought of other scientists buying it and using it on the basis of the original paper made me feel very guilty that we had not published our findings.
  4. It kept getting easier and easier to publish “negative results”. Journals such as Biology Open from Company of Biologists or PLoS ONE and preprint servers (see below) make this very easy.

Finally, it was a twitter conversation with Jim Woodgett convinced me that, when I had the time, I would write it up.

To which, he replied:

I added an acknowledgement to him in our paper! So that, together with the launch of bioRxiv, convinced me to get the paper online.

The Preprinting Experience

This paper was our first proper preprint. We had put an accepted version of our eLife paper on bioRxiv before it came out in print at eLife, but that doesn’t really count. For full disclosure, I am an affiliate of bioRxiv.

The preprint went up on 13th February and we submitted it straight to Biology Open the next day. I had to check with the Journal that it was OK to submit a deposited paper. At the time they didn’t have a preprint policy (although I knew that David Stephens had submitted his preprinted paper there and he told me their policy was about to change). Biology Open now accept preprinted papers – you can check which journals do and which ones don’t here.

My idea was that I just wanted to get the information into the public domain as fast as possible. The upshot was, I wasn’t so bothered about getting feedback on the manuscript. For those that don’t know: the idea is that you deposit your paper, get feedback, improve your paper then submit it for publication. In the end I did get some feedback via email (not on the bioRxiv comments section), and I was able to incorporate those changes into the revised version. I think next time, I’ll deposit the paper and wait one week while soliciting comments and then submit to a journal.

It was viewed quite a few times in the time while the paper was being considered by Biology Open. I spoke to a PI who told me that they had found the paper and stopped using pitstop as a result. I think this means getting the work out there was worth it after all.

Now it is out “properly” in Biology Open and anyone can read it.

Verdict: I was really impressed by Biology Open. The reviewing and editorial work were handled very fast. I guess it helps that the paper was very short, but it was very uncomplicated. I wanted to publish with Biology Open rather than PLoS ONE as the Company of Biologists support cell biology in the UK. Disclaimer: I am on the committee of the British Society of Cell Biology which receives funding from CoB.

Depositing the preprint at bioRxiv was easy and for this type of paper, it is a no-brainer. I’m still not sure to what extent we will preprint our work in the future. This is unchartered territory that is evolving all the time, we’ll see. I can say that the experience for this paper was 100% positive.


Dutta, D., Williamson, C. D., Cole, N. B. and Donaldson, J. G. (2012) Pitstop 2 is a potent inhibitor of clathrin-independent endocytosis. PLoS One 7, e45799.

Lemmon, S. K. and Traub, L. M. (2012) Getting in Touch with the Clathrin Terminal Domain. Traffic, 13, 511-9.

Stahlschmidt, W., Robertson, M. J., Robinson, P. J., McCluskey, A. and Haucke, V. (2014) Clathrin terminal domain-ligand interactions regulate sorting of mannose 6-phosphate receptors mediated by AP-1 and GGA adaptors. J Biol Chem. 289, 4906-18.

von Kleist, L., Stahlschmidt, W., Bulut, H., Gromova, K., Puchkov, D., Robertson, M. J., MacGregor, K. A., Tomilin, N., Pechstein, A., Chau, N. et al. (2011) Role of the clathrin terminal domain in regulating coated pit dynamics revealed by small molecule inhibition. Cell 146, 471-84.

Willox, A.K., Sahraoui, Y.M.E. & Royle, S.J. (2014) Non-specificity of Pitstop 2 in clathrin-mediated endocytosis Biol Open, doi: 10.1242/​bio.20147955.

Willox, A.K., Sahraoui, Y.M.E. & Royle, S.J. (2014) Non-specificity of Pitstop 2 in clathrin-mediated endocytosis bioRxiv, doi: 10.1101/002675.

The post title is taken from ‘Into The Great Wide Open’ by Tom Petty and The Heartbreakers from the LP of the same name.

Give, Give, Give Me More, More, More

A recent opinion piece published in eLife bemoaned the way that citations are used to judge academics because we are not even certain of the veracity of this information. The main complaint was that Google Scholar – a service that aggregates citations to articles using a computer program – may be less-than-reliable.

There are three main sources of citation statistics: Scopus, Web of Knowledge/Science and Google Scholar; although other sources are out there. These are commonly used and I checked out how comparable these databases are for articles from our lab.

The ratio of citations is approximately 1:1:1.2 for Scopus:WoK:GS. So Google Scholar is a bit like a footballer, it gives 120%.

I first did this comparison in 2012 and again in 2013. The ratio has remained constant, although these are the same articles, and it is a very limited dataset. In the eLife opinion piece, Eve Marder noted an extra ~30% citations for GS (although I calculated it as ~40%, 894/636=1.41). Talking to colleagues, they have also noticed this. It’s clear that there is some inflation with GS, although the degree of inflation may vary by field. So where do these extra citations come from?

  1. Future citations: GS is faster than Scopus and WoK. Articles appear there a few days after they are published, whereas it takes several weeks or months for the same articles to appear in Scopus and WoK.
  2. Other papers: some journals are not in Scopus and WoK. Again, these might be new journals that aren’t yet included at the others, but GS doesn’t discriminate and includes all papers it finds. One of our own papers (an invited review at a nascent OA journal) is not covered by Scopus and WoK*. GS picks up preprints whereas the others do not.
  3. Other stuff: GS picks up patents and PhD theses. While these are not traditional papers, published in traditional journals, they are clearly useful and should be aggregated.
  4. Garbage: GS does pick up some stuff that is not a real publication. One example is a product insert for an antibody, which has a reference section. Another is duplicate publications. It is quite good at spotting these and folding them into a single publication, but some slip through.

OK, Number 4 is worrying, but the other citations that GS detects versus Scopus and WoK are surely a good thing. I agree with the sentiment expressed in the eLife paper that we should be careful about what these numbers mean, but I don’t think we should just disregard citation statistics as suggested.

GS is free, while the others are subscription-based services. It did look for a while like Google was going to ditch Scholar, but a recent interview with the GS team (sorry, I can’t find the link) suggests that they are going to keep it active and possibly develop it further. Checking out your citations is not just an ego-trip, it’s a good way to find out about articles that are related to your own work. GS has a nice feature that send you an email whenever it detects a citation for your profile. The downside of GS is that its terms of service do not permit scraping and reuse, whereas downloading of subsets of the other databases is allowed.

In summary, I am a fan of Google Scholar. My page is here.


* = I looked into this a bit more and the paper is actually in WoK, it has no Title and it has 7 citations (versus 12 in GS). Although it doesn’t come up in a search for Fiona or for me.



However, I know from GS that this paper was also cited in a paper by the Cancer Genome Atlas Network in Nature. WoK listed this paper as having 0 references and 0 citations(!). Does any of this matter? Well, yes. WoK is a Thomson Reuters product and is used as the basis for their dreaded Impact Factor – which (like it or not) is still widely used for decision making. Also many Universities use WoK information in their hiring and promotions processes.

The post title comes from ‘Give, Give, Give Me More, More, More’ by The Wonder Stuff from the LP ‘Eight Legged Groove Machine’. Finding a post title was difficult this time. I passed on: Pigs (Three Different Ones) and Juxtapozed with U. My iTunes library is lacking songs about citations…

A Day In The Life

#paperoftheday #potd

A common complaint from other PIs is that they “don’t read enough any more”. I feel like this too and a solution was proposed by a friend of a friend*: try to read one paper per day.

This seemed like a good idea and I started to do this in 2013. The rules, obviously, can be set by you. Here’s my version:

  1. Read one paper each working day.
  2. If I am away, or reviewing a paper for a journal or colleague, then I get a pass.
  3. Read it sufficiently to be able to explain it to somebody else, i.e. don’t just scan the abstract and look at the figures. Really read it and understand it. Scan and skim as many other papers as you normally would!
  4. Only papers reporting primary research count towards #paperoftheday.
  5. If it was really good or worth telling people about – tweet about it.
  6. Make a simple database in Excel or Papers – this helps you keep track, make notes about the paper (to see if you meet #3) and allows you to find the paper easily in the future (this last point turned out to be very useful).

I started this in 2013 (for one full year) and am trying to continue in 2014. I feel that this is succeeding in making me read more than I would have otherwise done.

My stats for 2013 were:

  • 85% success rate. Filling that last 15% will be tough.
  • Stats errors in 48% of papers! Most common error was incorrect use of Student’s  t-test.
  • 68% of papers were from 2013 and 22% were from 2009-2012.

The big surprise was which journals I read most:

  1. J Cell Biol 13
  2. PLOS One 12
  3. Nat Cell Biol 10
  4. PNAS 10
  5. Curr Biol 9
  6. Mol Biol Cell 8
  7. Nature 8
  8. Dev Cell 7
  9. eLife 7
  10. Nature Methods 7
  11. Cell 6
  12. Neuron 6
  13. Traffic 6
  14. J Cell Sci 4
  15. Science 4

I thought that Cell would be much higher and PNAS would be much lower. Since where we publish is dictated by who is likely to see and read the paper, this list was thought-provoking.

*I think this was a colleague of @david_s_bristol who suggested it, sometime in 2012.

The post title is of course from A Day in The Life – The Beatles from the LP Sgt Pepper’s Lonely Hearts Club Band. For the first line…

So Long

How long does it take to publish a paper?

I posted the picture below on Twitter to show how long it takes for us to publish a paper.


The answer is 235 days. This is the median time from submission at the first journal to publication online or in print. The data are from our last ten papers.

The infographic proved popular with 40 retweets and 22 favourites. It was pointed out to me that the a few things would improve this visualisation:

1. Showing the names of the journals

2. Showing when the 1st submission was relative to the 1st submission at the journal that finally accepted the paper

3. What about reviews and other types of publication.

I am working on updating the graph to show all of these things… watch this space.

My point was really to show (perhaps to non-scientists) how long the process of publishing a paper can be. There is other information that can be gleaned from this, e.g. what proportion of time is at the journal’s side and how much is at our end?

The people who are eager to see which journals perform badly (slowly) will be disappointed: this is a very small subset of papers from one lab. I’d be interested in scraping the information on journal tardiness on a larger scale and synthesising this so that it can inform journal choice. Recently though major publishers have taken steps to make this information less accessible so don’t hold your breath.

The title of this post is from So Long by Cian Ciarán from the LP ‘Outside In’